What has changed
Anthropic launched Claude Opus 5 on 24 July 2026. The company’s positioning is clear: it does not present it merely as a fast-response model, but as one designed to solve programming and knowledge tasks that require reviewing results, using tools and maintaining context across several steps. The model is available through the API under the claude-opus-5 identifier and, according to the current documentation, is listed as active. The earliest possible retirement date shown is 24 July 2027; this is not a confirmed retirement date, but a minimum boundary within Anthropic’s lifecycle policy.
The main commercial change is API pricing: 5 dollars per million input tokens and 25 dollars per million output tokens, the same base rates that Anthropic assigns to Opus 4.8. This is accompanied by an “effort” parameter that lets users decide how much reasoning and token consumption to devote to each request. The implicit promise is that lower effort can contain spending and latency, while higher levels seek to prioritize quality for more difficult tasks.
The technical sheet places its context window at one million tokens and its usual maximum output at 128,000 tokens; the batch API, still marked beta, can reach 300,000 output tokens. The model accepts text and images as input and produces text. Its reliable knowledge and training-data cutoff is May 2026. It therefore should not be treated as a self-contained source of current information on events after that date: such cases require access to external sources, search tools or documentation supplied by the organization using it.
There are also operational changes that affect existing integrations. Anthropic’s documentation indicates that parameters such as temperature, top_p and top_k are deprecated for Claude 4.7 models and later; if configured outside their default values, they may produce an error. For teams upgrading an application, migration is not simply a matter of changing the model name: they need to review parameters, regression tests, token budgets and oversight rules.
Capabilities and evidence: what the provider claims
Anthropic says Opus 5 achieves leading results in some programming and knowledge-work evaluations. The tests it highlights include Frontier-Bench v0.1, CursorBench 3.2, ARC-AGI 3, Zapier AutomationBench and OSWorld 2.0. In programming, it says it exceeds the other models tested in Frontier-Bench and, at maximum effort, comes within less than half a percentage point of Claude Fable 5’s top result in CursorBench, at half the cost per task. In automation and computer use, the company says it delivers a better results-to-cost balance than other models included in its charts.
The company also reports improvements over Opus 4.8 in internal life-sciences evaluations, including organic chemistry and bioinformatics tasks. Specifically, it reports a 10.2-percentage-point difference in an internal benchmark for inferring molecular structures from spectroscopy and 7.7 points in a task related to protein variants. These figures describe Anthropic’s measurements and do not, by themselves, amount to clinical, scientific or regulatory validation of the model’s use in applied research.
The published evidence combines several kinds of material. There are benchmarks with named tests and summarized results; tests run by Anthropic; and testimonials from companies involved in early access. Testimonials can provide examples of adoption, but they are not independent comparisons and do not replace a reproducible assessment under each organization’s conditions. In addition, the launch note specifies that at least one Frontier-Bench measurement was run in an internal environment, with five attempts per task and with Opus 4.8 as a fallback when safety classifiers rejected requests to Opus 5 or Fable 5.
That last detail matters when interpreting the charts. A score obtained in a system involving retries, tools, a particular effort setting or a fallback model should not be read as the isolated ability of a single model in a single call. Nor is it safe to transfer a benchmark ranking directly to a real business process. Production involves factors that many tests do not fully capture: context quality, tool permissions, data state, instruction format, error tolerance, traceability requirements and the cost of human review.
Anthropic’s prompting guide offers a practical signal of how it expects the model to be used. It recommends structuring complex instructions with XML tags, providing relevant and diverse examples, and carefully ordering long documents. For inputs above 20,000 tokens, it advises putting extensive information first and the query afterwards; it also recommends requesting direct quotations from supplied documents before performing analysis. These practices can reduce ambiguity, but they do not guarantee factual accuracy or remove the need for external controls.
Documented limits and risks
The launch does not remove safety restrictions. Anthropic says Opus 5 does not cross the dual-use risk capability threshold in biology and cybersecurity, and that it remains behind Mythos 5 in biological research and offensive cybersecurity. Even so, it acknowledges meaningful gains in identifying software vulnerabilities as a consequence of its general capability. It distinguishes detection from exploit creation: according to its OSS-Fuzz assessment, Opus 5 approaches Mythos 5 at finding vulnerabilities, but remains substantially behind it at developing exploits.
To manage that risk, Anthropic applies classifiers to certain cybersecurity requests. The policy described allows source-code vulnerability discovery, but blocks binary-based vulnerability analysis, penetration testing and exploit generation. When a request is flagged in Claude.ai, Claude Code or Claude Cowork, the default behavior is to redirect it to Opus 4.8. The API can likewise enable automatic fallbacks. This architecture limits some uses, but raises another governance issue: an organization should record which model ultimately answered, because behavior, cost and restrictions may differ from those initially expected.
Anthropic also recognizes important limitations in long-duration autonomous biological research, precisely the area it identifies as most relevant to biological risks. The existence of safeguards and safety testing does not mean the system is infallible. Automated controls can create unjustified blocks, incorrect classifications or responses from an alternative model. Conversely, a request that does not trigger a classifier is not thereby validated as correct, safe or suitable for a high-impact decision.
For information tasks, familiar language-model risks remain: plausible but incorrect answers, nonexistent or wrongly attributed citations, omission of relevant conditions and mistakes when interpreting data. The May 2026 knowledge cutoff adds a specific temporal limitation. For legal, medical, financial, scientific or security analysis, Opus 5 should operate as support under expert review and with verification against primary sources, not as a replacement for professional judgment.
There is also a cost of technological dependence. Anthropic’s documentation warns that models can be active, legacy, deprecated or retired, and that requests to retired models will fail. Although the company says it will give customers with active deployments at least 60 days’ notice before retiring public models, applications should be designed with compatibility testing, provider abstraction, version-level metrics and rollback plans. Service continuity should not depend on a particular model retaining the same behavior indefinitely.
Practical impact: deciding whether it fits
For development teams, Opus 5 is most relevant when work requires understanding a repository, investigating a bug, proposing coordinated changes, running tools and validating results. Its potential value is not measured by the wording of a brief response, but by the total cost of closing a task: tokens, tool calls, latency, human review time, avoided failures and repeated runs. A useful test should use representative issues, modules, documentation and acceptance criteria, while avoiding unnecessary exposure of secrets or personal data.
In knowledge work, the model may help summarize extensive files, compare documents, extract table structures, prepare drafts or identify issues for review. The more cautious approach is to require it to distinguish evidence, inferences and recommendations; request references to the supplied material; and ensure that output is not used to execute external actions without confirmation. The ability to retain context does not by itself solve a data-quality problem: if documents are contradictory, incomplete or incorrectly dated, a well-written synthesis can still be wrong.
The effort setting makes it possible to experiment with a portfolio of tasks. Routine requests can be assessed at lower levels, whereas complex debugging, numerical reasoning or high-impact document analysis may justify higher effort and additional controls. Yet the cheapest option does not always minimize final cost, and the most intensive one does not always provide the best answer. It is sensible to measure success rate, input and output tokens, response time, human interventions, critical errors and variation across runs. Comparisons should be made against the current process, not only against another model.
The technical guide itself suggests avoiding legacy instructions that demand excessive self-checking when migrating to Opus 5, because they may add tokens and latency. That advice does not mean eliminating external review. An internal instruction asking the model to revisit its work is very different from an independent control: automated tests, data validation, code review, dual approval or human authorization before an irreversible operation.
An adoption decision should include a bounded pilot. It is advisable to select tasks with verifiable outcomes, define a quality threshold and budget, test different effort levels, record when filters or fallbacks intervene, and review samples of failures. If the model uses tools able to write, deploy, purchase, change permissions or send communications, permissions should be minimal and reversible. Autonomy should grow gradually, and only after reliability has been demonstrated in a controlled environment.
Conclusions
According to Anthropic’s information, Claude Opus 5 is an update focused on making advanced-model use more viable for programming, automation and knowledge analysis without increasing the base price compared with Opus 4.8. Its one-million-token context window, effort control and focus on multi-step workflows describe a product intended for broader integrations than a conventional chat interface.
However, performance claims must be read in context. They come largely from the provider and depend on benchmarks, evaluation harnesses, effort levels and, in certain cases, fallback mechanisms. The results are relevant signals, but not a universal demonstration of superiority or a guarantee of performance in a specific environment. The decisive question for an organization is not whether Opus 5 leads a table, but whether it measurably improves a particular process with acceptable risks and costs.
The model also does not eliminate issues involving safety, factual freshness, hallucinations or version dependence. Anthropic documents cybersecurity restrictions, limits in prolonged autonomous research and a knowledge cutoff date. Used with well-governed data, local testing, observability and human review proportionate to risk, it can support complex tasks. Used as an autonomous authority or as a substitute for controls, its limitations remain decisive.
Open questions
- No independent evaluation has been conducted of the benchmarks, charts or customer testimonials cited by Anthropic in its launch presentation.
- Cost-per-task metrics depend on effort settings, available tools, retries, context length and the price of human oversight; they cannot automatically be extrapolated to every case.
- The documentation reviewed gives a minimum retirement date rather than a definitive date, and partner-operated platforms may apply different lifecycle schedules.
- The documentation does not make it possible to infer the error rate in a specific domain, language, repository or document set without local testing.
- Classifiers and automated fallbacks may change the final answer compared with a direct Opus 5 run; the actual incidence depends on configuration and request type.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction