Laboratory
Entity that develops or publishes the model and maintains its technical and safety documentation.
Explore the documented models, published evaluations and primary sources associated with this organisation.
Model oriented to code and knowledge work, with controls that allow effort to be adjusted to each task.
16 SEP 2026↗ LONG-RUNNING REASONINGLong-running reasoning and agents.
16 SEP 2026↗ COST/CAPABILITY BALANCECoding, research, and professional tool-based workflows.
16 SEP 2026↗ HIGH EFFICIENCYHigh volume, subagents, and structured extraction.
16 SEP 2026↗ PREVIOUS FRONTIER · AGENTSPrevious Opus version focused on code, agents, and computer use, retained for migration and historical comparison.
18 SEP 2026↗ AGENTS AND COMPUTER USEPrevious Sonnet generation for code, agents, and interface automation, with official computer use results.
18 SEP 2026↗ AGENTIC CODING AND KNOWLEDGE WORKNext-generation Opus for agentic coding, computer use, and extended professional work, with a one-million-token context.
29 SEP 2026↗ BALANCED INTELLIGENCE AND SPEEDSonnet update for everyday tasks, code, and documents, with improved speed and a million tokens of context.
29 SEP 2026↗Figures are shown with the context reported by their source. A provider result is not an independent comparison and does not replace your own evaluation.
| Model | Benchmark | Result | Metric |
|---|---|---|---|
| Claude Opus 5 | Terminal-Bench Science 0.1 | 29,0 % | Accuracy |
| Claude Opus 5 | CursorBench 3.2 | 70,0 % | Accuracy |
| Claude Opus 5 | AutomationBench | 26,9 % | Accuracy |
| Claude Fable 5.1 | Terminal-Bench Science 0.1 | 52,6 % | Accuracy |
| Claude Fable 5.1 | CursorBench 3.2 | 73,4 % | Accuracy |
| Claude Fable 5.1 | Humanity's Last Exam | 65,0 % | With tools |
| Claude Sonnet 5 | OSWorld-Verified | Curve by effort | Success and cost |
| Claude Sonnet 5 | BrowseComp | Curve by effort | Accuracy and cost |
| Claude Haiku 4.5 | SWE-bench Verified | 73,3 % | Resolved |
| Claude Haiku 4.5 | Terminal-Bench | 41,75 % | With reasoning |
| Claude Opus 4.5 | Evaluaciones Claude Opus 4.5 | Published | Code, agents, and computer |
| Claude Sonnet 4.5 | OSWorld | 61,4 % | Computer tasks solved |
| Claude Opus 5.5 | Terminal-Bench 4.0 | 66,4 % | Accuracy |
| Claude Opus 5.5 | GDPval-AA v2.1 | 1.846 | Elo |
| Claude Sonnet 5.5 | Terminal-Bench 4.0 | 70,6 % | Accuracy |
An organization, a product, and a model are not the same unit. Inferama separates them to avoid attributing capabilities or commercial terms to the wrong item.
Entity that develops or publishes the model and maintains its technical and safety documentation.
Identifiable version with limits, modalities, and behavior that may change between releases.
API, application, associated cloud, or commercial plan; each channel may have different pricing, retention, and limits.
Every claim must retain the official page consulted and the verification date.
A template for estimating tool-workflow costs by round, separating input and output tokens from additional charges, and checking the calculation against usage logs.
28 Sep 2026 ↗ GUIAA price per million tokens is not enough to budget a multi-step workflow. Learn how to add up input, output, cache, retries, and review, and set limits before you run it.
25 Sep 2026 ↗ ANALISISAnthropic says the same text may produce more tokens with Sonnet 5 than with Sonnet 4.6. A test using your own corpus can show whether that reduces usable context or changes cost per task in a specific integration.
23 Sep 2026 ↗ ANALISISSonnet 4.5’s scores on biological research tasks justify internal evaluation, not a conclusion about its experimental reliability. What Protocol QA, LAB-Bench, and BixBench measure, what they do not, and how to design a controlled retrospective test.
23 Sep 2026 ↗ GUIAThe price per call does not, by itself, show how much an accepted response costs. This guide provides a reproducible formula for adding input and output tokens, retries, and review, with three illustrative scenarios for Claude Haiku 4.5.
23 Sep 2026 ↗ ANALISISA benchmark score does not prove that a model can operate an application safely. We propose a bounded evaluation to measure success, errors, recovery, latency, and the need for supervision.
23 Sep 2026 ↗