Exact version
Check the model identifier, date and whether the alias can change.
Published results and conditions declared by their sources. Figures are comparable only when version, configuration, metric and date match.
| Model | Organization | Result | Metric | Conditions |
|---|---|---|---|---|
| GPT‑6 Astra | OpenAI | 96,0 % | Accuracy | Provider result. |
| DeepSeek V4.1 Flash | DeepSeek | 90,9 % | Accuracy | DeepSeek Harness, max effort, and published configuration. |
| Qwen3‑Max Thinking | Alibaba Qwen | 92,8 % | Accuracy | Inference scaling published by Qwen. |
Check the model identifier, date and whether the alias can change.
Review tools, effort, number of attempts, prompt and inference budget.
Validate the result with tasks, languages, formats and errors representative of your product.
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
A moving alias and a snapshot may produce different results. Record the model, date, and provider.
Prompt, reasoning, tools, number of attempts, and budget must be equivalent for comparison.
A small difference may disappear between runs. Keep samples, dispersion, and failures, not just the average.
Validate whether the improvement holds across languages, formats, and real cases close to your product.
A metric is only useful when we know the task, the conditions, and the error that matters. The result begins with the design of the test.
A useful comparison between DeepSeek R1 and DeepSeek V3.2 does not begin with a published benchmark, but with an organization’s own task, a frozen corpus, and an explicit definition of material error. This protocol makes it possible to measure whether R1’s additional deliberation reduces errors and human review, or whether V3.2 reaches the same threshold with lower latency, consumption, and complexity.
22 Sep 2026 ↗ ANALISISAn AIME 2024 figure can look straightforward while summarizing very different protocols. This guide explains what the dataset represents, how results change with the number of attempts, answer-selection methods, tools, and evaluation harnesses, and what information to request before comparing models.
22 Sep 2026 ↗