Exact version
Check the model identifier, date and whether the alias can change.
Published results and conditions declared by their sources. Figures are comparable only when version, configuration, metric and date match.
| Model | Organization | Result | Metric | Conditions |
|---|---|---|---|---|
| GPT‑5.6 Sol | OpenAI | 52,7 % | Accuracy | Official GPT‑5.6 table. |
| GPT‑5.6 Terra | OpenAI | 50,4 % | Accuracy | Official GPT‑5.6 table. |
| GPT‑5.6 Luna | OpenAI | 50,3 % | Accuracy | Official GPT‑5.6 table. |
Check the model identifier, date and whether the alias can change.
Review tools, effort, number of attempts, prompt and inference budget.
Validate the result with tasks, languages, formats and errors representative of your product.
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
A moving alias and a snapshot may produce different results. Record the model, date, and provider.
Prompt, reasoning, tools, number of attempts, and budget must be equivalent for comparison.
A small difference may disappear between runs. Keep samples, dispersion, and failures, not just the average.
Validate whether the improvement holds across languages, formats, and real cases close to your product.
A metric is only useful when we know the task, the conditions, and the error that matters. The result begins with the design of the test.
AutomationBench evaluates whether an agent can complete workflows across simulated SaaS applications and bring them to a verifiable final state. This guide explains what that evidence means, how to read its metrics, and what additional testing is needed before extrapolating a result to a real company.
22 Sep 2026 ↗ PAPERAgents’ Last Exam evaluates agents on professional tasks within operating-system environments and with verifiable outcomes. Its design is useful for studying workflow execution, but an aggregate success figure does not demonstrate that a job can be automated. Interpreting it requires details about the tasks, environment, harness, model, tools, budget, and evaluated version.
22 Sep 2026 ↗