A new frontier model, with caveats about what “better” means
Anthropic presents Claude Opus 5.5 as a new version of its Opus model family. The announcement brings together three claims: competitive results on evaluations, a lower cost than Opus 5, and stronger safeguards for sensitive capabilities. These claims matter to teams comparing models, but they do not all have the same kind of backing. Some describe product specifications; others are results from Anthropic evaluations or comparisons reported by the media.
The company calls it the best-performing model it has tested to date. That wording should be read as a statement from Anthropic, not as an independent conclusion that applies to every task. Results can vary depending on the evaluation set, effort setting, permitted tools and token budget. The sources provided do not offer a complete comparison that holds all of these conditions constant.
It is also important to distinguish “Opus 5.5 costs less per token” from “a task costs less.” The cost of a run depends, among other things, on how much text it processes and generates, whether caching is used, and how many calls or tools the workflow requires. A lower unit price does not, on its own, guarantee an equivalent reduction in the final cost.
What has been reported about performance and efficiency
Xataka’s coverage attributes a score of 1846 Elo on GDPval-AA v2.1 to Anthropic. The evaluation focuses on professional-work tasks. In the same comparison, the company places Fable 5.1 at 1735 and Opus 5 at 1708. Anthropic is also reported to say that Opus 5.5, at medium effort, outperforms GPT-6 Astra configured at maximum effort on this test, at approximately one-fifth of the cost per task.
The score is an indication of performance on the cited evaluation, not a universal measure of quality. Artificial Analysis maintains the benchmark table and presents results alongside intervals and effort configurations. Consulting that table can help explain the context behind a score, but it does not, by itself, establish per-task costs or validate every condition used by the provider.
Secondary coverage also refers to execution-cost reductions of up to 40%. That figure does not necessarily mean a 40% reduction in every per-token price. Quartz, meanwhile, reports rates of $4 per million input tokens and $20 per million output tokens, compared with $5 and $25 for Opus 5. On those figures, the standard rates are 20% lower; larger per-task savings are a different metric and depend on usage.
Anthropic also attributes higher speed and lower token use on some workloads to the model. The information provided does not describe an independent protocol that would allow those results to be applied to every use case or quantify how response times change under each configuration. An operational comparison needs to measure quality, total cost and latency on representative tasks.
How to interpret the reported figures
These metrics describe different things and should not be treated as interchangeable.
| Figure | What it indicates | What it does not prove on its own |
|---|---|---|
| GDPval-AA v2.1 Elo score | Relative performance within that evaluation and configuration. | General superiority across all professional tasks. |
| Price per million tokens | The unit cost of input or output under the published rate. | The final cost of a task or a complete agent workflow. |
| Announced savings per task | A cost estimate for a task under the reported evaluation method. | The same savings in every application, workload or system. |
| Speed or lower token use | Efficiency reported for certain tests. | A consistent latency improvement under all conditions. |
Pricing, context and availability
The official model documentation specifies a one-million-token context window and a maximum output of 128,000 tokens. These limits are relevant when analysing long documents or generating lengthy responses, but they do not guarantee that a request of that size will be suitable, fast or inexpensive. Usable context and effective limits may also depend on the product or platform used to access the model.
Quartz reports prices of $4 per million input tokens and $20 per million output tokens. It compares these rates with Opus 5, at $5 per million input tokens and $25 per million output tokens. The official model documentation is the most appropriate reference for verifying current prices; the secondary coverage provides the explicit numerical comparison reproduced here.
Anthropic says it has increased usage limits on its subscription plans. The information available in the supplied sources does not specify the exact scope of each increase, the quotas that apply to each plan, or whether conditions are the same across access channels. Anyone evaluating the model should therefore check the terms for their specific plan and platform rather than assume that one limit applies everywhere.
The documentation lists platforms and model identifiers, but the information retrieved is not sufficient to detail every channel, date or availability condition here without risking an error. Nor does it provide a complete, up-to-date comparison with Fable 5.1. Media coverage focuses mainly on selected results and claims attributed to Anthropic.
A practical check before budgeting
- 01Check the price and model identifier in the official documentation and on the platform you plan to use.
- 02Separate input, output and cache costs, then estimate the application’s actual token volume.
- 03Confirm the usage limits for the specific plan and how long requests are handled.
- 04Measure cost, latency and quality on a sample of your own tasks before projecting savings.
Safeguards and limits on access to capabilities
Anthropic announces stronger safeguards alongside the model. Xataka’s coverage describes restrictions on what Opus 5.5 can do and highlights controls associated with sensitive capabilities, particularly in cybersecurity and biology. However, the supplied material does not specify in enough detail which requests are blocked, what criteria are applied, or whether restrictions vary by user, plan or environment.
It is therefore not possible to conclude from these sources that a specific access regime is reserved for particular groups, or to describe exceptions or approval procedures. The available documentation supports the claim that Anthropic presents the deployment as having safeguards; secondary coverage reports restrictions, but does not provide a complete policy from which their application in every case can be reconstructed.
For an organisation considering the model, the question is not only whether a capability is available, but also how it responds to legitimate security or research tasks, what controls are in place, and what review options the provider offers. Those conditions require consulting the security documentation and conducting authorised tests in the intended environment. They should not be inferred solely from a benchmark score or a headline.
What a team should test before migrating
The published figures make Opus 5.5 worth evaluating, but they are not enough to recommend a migration. A comparison should use the same set of tasks, instructions, tools, output limits and budget. If one model can use more effort or make more calls, the difference in quality may reflect the configuration as well as the model itself.
It is also useful to record task completion rates, errors requiring human intervention, total time and the cost of each run separately. For agent workflows, include tool calls and repeated context tokens as well as generation charges. A limited trial using representative cases and criteria established in advance is more useful than extrapolating from a single ranking.
The final decision will depend on the application. A lower price may matter for a high-volume system; a larger context window may help with lengthy documents; and safety restrictions may be decisive for specialised tasks. However, the supplied sources do not provide enough independent evidence to conclude that Opus 5.5 is the most efficient option in all of these scenarios. Missing information should be treated as uncertainty, not as a confirmed advantage or disadvantage.
Open questions
- No complete independent evaluation is provided that reproduces Anthropic’s reported performance comparisons and per-task cost claims.
- The GDPval-AA v2.1 score should be checked against the evaluation owner’s table and configurations; it does not establish general superiority.
- The exact conditions used to compare Opus 5.5 at medium effort with GPT-6 Astra at maximum effort are not established.
- Explicit rates are reported in secondary coverage; they should be verified in the official documentation and on the relevant platform before budgeting.
- The new usage limits, their distribution by plan and availability across all channels are not fully detailed.
- The sources do not provide a comprehensive official description of safety restrictions, affected tasks or possible differences in access.
- The sources do not establish a complete, up-to-date comparison of Opus 5.5, Opus 5 and Fable 5.1 beyond the specific metrics cited.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction