Quality is measured 2 main ways:
1) End-to-end: User query -> to task resolution. These are aider style benchmarks answering the question of actual task completion
2) Apply Quality: Syntax correctness, character diff, etc..
The error rate for large vs fast is around 2%. If you're doing code edits that are extremely complex or on obscure languages - large is the better option. There's also an auto option to route to the model we think is best for a task