Exact version
Check the model identifier, date and whether the alias can change.
Published results and conditions declared by their sources. Figures are comparable only when version, configuration, metric and date match.
| Model | Organization | Result | Metric | Conditions |
|---|---|---|---|---|
| Claude Sonnet 5 | Anthropic | Curve by effort | Accuracy and cost | Methodology reviewed on 30 JUN 2026. |
Check the model identifier, date and whether the alias can change.
Review tools, effort, number of attempts, prompt and inference budget.
Validate the result with tasks, languages, formats and errors representative of your product.
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
A moving alias and a snapshot may produce different results. Record the model, date, and provider.
Prompt, reasoning, tools, number of attempts, and budget must be equivalent for comparison.
A small difference may disappear between runs. Keep samples, dispersion, and failures, not just the average.
Validate whether the improvement holds across languages, formats, and real cases close to your product.
A metric is only useful when we know the task, the conditions, and the error that matters. The result begins with the design of the test.