Ilustración editorial para Claude Opus 4.5 vs Claude Sonnet 4.5: cuándo pagar más para corregir incidencias de código y cuándo no
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The decision is not which model seems better, but which one resolves an issue at the lowest total cost

Comparing Claude Opus 4.5 and Claude Sonnet 4.5 for software maintenance requires narrowing the question. Token pricing, an isolated demonstration, or a benchmark result alone do not tell a team which option is the better fit. The decision unit should be the correctly resolved and accepted issue: a change that reproduces the failure context, passes target and regression tests, does not improperly expand the scope, and requires a reasonable amount of human review.

In this comparison, a price premium for Opus 4.5 would be justified only if it translates into a measurable operational outcome. That may mean a higher full-resolution rate, fewer regressions, fewer engineer iterations, or less time before an acceptable patch is available. If Sonnet 4.5 reaches the same acceptance thresholds at a lower cost, paying more does not necessarily create value for that kind of ticket.

The analysis must separate inference cost from labor cost. A low-cost model that produces incomplete patches, requires its diagnosis to be debugged, or makes out-of-scope changes can become more expensive after retries, tool execution, and review are included. Conversely, a higher-priced model should not receive credit for a patch that looks plausible but does not pass a representative test suite.

This article does not declare a general winner. It proposes a local test that lets engineering leaders, AI platform teams, and buyers make a decision using their own evidence. It is also compatible with the comparisons index, the Claude Opus 4.5 and Claude Sonnet 4.5 model profiles, and Anthropic’s information about the models and their access channels.

02

Prepare a frozen, reproducible corpus before running the models

The test begins with the corpus, not the prompt. Every issue should be associated with a repository, an immutable base commit, an execution environment, and a verifiable definition of the failure. Keep the original ticket, but write a task version that removes information created after the point in time you want to simulate, such as a link to the already-known fix or comments that reveal the solution.

For every case, create a minimal reproduction that fails on the base commit and a suite that can verify the fix. The evaluation must distinguish a cosmetic modification from a functional repair. A strong practice is to define tests that should have failed before the fix and tests that already passed and must continue to pass. This approach aligns with the validation logic used to distinguish resolution from regression in SWE-bench Verified, although results from an internal corpus must not be combined with that benchmark’s results.

SWE-bench is a useful methodological reference, not a substitute for a local corpus. The original work describes a collection of issues extracted from real Python repositories; that helps explain the demands of a repository-editing task. However, its language distribution, dependencies, issue age, and evaluation rules may differ from those of a given team. An internal result should be published as an internal result.

Exclude tickets without a reasonably stable reproduction, changes that depend on uncontrolled external services, vulnerabilities requiring a specific disclosure process, and tasks whose acceptance criterion is purely subjective. Recording those exclusions prevents the final sample from appearing more representative than it actually is.

Freezing each issue

  1. 01Select a closed issue whose known solution will not be supplied to the model.
  2. 02Pin the repository, base commit, dependency versions, operating system, and test command.
  3. 03Verify that the reproduction fails on the base commit and retain the logs.
  4. 04Define target tests, regression tests, and a review rubric before running either model.
  5. 05Archive the ticket, evaluation scripts, and run artifacts under an internal identifier.
03

Keep conditions equivalent, but do not assume equivalent means identical

Record each model’s exact identifier, the execution date and time, access channel, region or endpoint, infrastructure provider, context configuration, and granted tools. For Opus 4.5, Anthropic’s documentation identifies the API model as claude-opus-4-5-20251101. Model lifecycle documentation should be checked at the start of every campaign to confirm that both identifiers remain active and to anticipate retirements.

The Sonnet 4.5 announcement documents API access and its launch input and output pricing. That information is not enough to calculate a final cost: billing can depend on the channel, cache use, region, tools, and current commercial terms. In Amazon Bedrock, for example, provider documentation describes differences between global and regional endpoints, as well as specific availability conditions. Do not compare a regional run of one model with a global run of the other without declaring the difference.

Use the same system prompt, task description, response format, working directory, read and write permissions, authorized commands, network access, and iteration limit. If one model has a mode or capability that the other does not have in the selected channel, do not treat it as a strictly equivalent test. You may measure it as a separate operational scenario, but it must be labeled accordingly.

Retry policies and limits must also be fixed. A retry after a transient infrastructure error may be reasonable; multiple restarts because the first result was poor change the intervention available. The rule must apply to both models, and every billable attempt must be counted.

Variables to record for every run

VariableControl ruleWhy it matters
Model and identifierPin the exact ID and consultation datePrevents comparison across different revisions or lifecycle stages
Channel, region, and endpointKeep them the same or separate cohortsThey can change availability, latency, and price
Tools and permissionsUse the same set and the same limitsThey affect inspection and validation capability
Context and iterationsUse the same maximum budget per ticketPrevents giving one model more opportunities
RetriesSet the policy in advance and log all of themIncludes the cost of failures and recoveries
Test environmentPin the image and dependenciesReduces outcomes caused by environment drift
04

Separate tickets by difficulty and failure mechanism

A global average can conceal the information that actually determines the purchase decision. Classify cases before running the models. A first stratum can contain localized fixes: an incorrect validation, an edge condition, or a data transformation limited to one or a few files. These are tasks in which a more economical model may reach the acceptance threshold quickly.

A second stratum should include failures that require tracing dependencies across modules. For example, an internal interface change that breaks serialization, validation, and remote consumers, or a defect where the symptom appears in a different layer from the cause. Here, it is reasonable to investigate whether greater planning or exploration capability reduces iterations, but the result must be measured rather than inferred from the model category.

The third stratum covers ambiguous or hard-to-reproduce issues. They may involve concurrency, shared state, specific configurations, or incomplete requirements. The goal is not to reward a lengthy explanation: it is to determine whether the model forms testable hypotheses, obtains evidence using the permitted tools, and limits the change to the most strongly supported cause.

Also label language, repository size, modified surface area, test type, and the presence of external dependencies. Those labels make it possible to determine whether a difference arises from real difficulty or from an accidental concentration of tickets in one language or module.

05

Define complete success before viewing the results

The success criterion should combine automated validation and human review. Classify a patch as a complete success only when it passes target tests, preserves relevant regression tests, and meets the scope rubric. If the repository has a broad test suite that is feasible within the budget, run it; if not, declare what coverage was omitted and why.

Human review must be blind to the model. Give reviewers the diff, the produced diagnosis, test results, and the ticket, but not the model name or cost. Ask for a defined decision: accept, accept with minor changes, reject for incomplete correctness, reject for regression, reject for excessive scope, or another category specified in advance.

It is useful to retain a partial-success category, but not to use it to inflate the resolution rate. A patch that correctly locates the affected component but misses an edge condition can be useful for investigating diagnosis quality. It is not equivalent to a resolved issue. Similarly, a correct fix that modifies unrelated files without justification may require enough intervention that it should not count as a complete success.

Review instructions should prohibit accepting changes that disable tests, weaken assertions, or introduce generic exceptions merely to make an error disappear. They should also state that a new test alone does not prove that the implementation is correct.

Minimum outcome rubric

ClassificationConditionUse in the decision
Complete successTarget and regression tests pass; review accepts the scopeCounts toward cost per resolved issue
Partial successVerifiable progress, but correctness or review is still lackingAnalyze separately; does not count as resolution
Technical failureDoes not reproduce, does not compile, fails tests, or regresses behaviorCounts consumed cost and failure cause
Scope failureChange is excessive, risky, or difficult to maintainCounts as not accepted; quantifies additional review
06

Measure costs, intervention, and end-to-end time

The main metric can be expressed as total batch cost divided by the number of accepted complete successes. In the numerator, include input, output, cache, and any billable concept applicable to the channel. Add retries, failed calls, tool use when it has a cost, and human review or correction time if the goal is to decide operational cost rather than API cost alone.

Also report the complete-success rate, partial-success rate, detected regressions, number of human interventions, and end-to-end latency. Latency is not necessarily model time: an agent may wait for tools, repeat tests, or block continuous integration resources. Record inference time, tool time, and human time separately whenever possible.

To evaluate diagnosis, use a simple rubric: identification of symptoms, causal hypothesis, evidence gathered, explanation of the modification, and known limitations. A diagnosis can be useful even in a failure, but it must be evaluated without confusing narrative quality with correctness. The score should have reference examples and, where there are multiple reviewers, a rule for resolving disagreements.

Report distributions and results by stratum, not just averages. A small number of complex issues can dominate average cost. Also repeat some or all of the batch: run-to-run variability can change the conclusion when the difference between models is small.

Operational calculation per ticket

  1. 01Add all billed amounts and tool costs attributed to the ticket.
  2. 02Record human review time and apply an internal rate defined before the test.
  3. 03Label the outcome using the blind rubric and retain test logs.
  4. 04Include the costs of accepted and unaccepted cases in the batch calculation.
  5. 05Divide total cost by complete successes; also publish the acceptance rate and its variation by stratum.
07

Interpret the Opus 4.5 premium through thresholds, not model prestige

Opus 4.5 would justify a premium in a specific stratum if its improvement in complete resolutions offsets its additional costs and the review it avoids. This could occur in issues involving relationships across modules, ambiguous diagnoses, or costly test cycles, but it is a hypothesis the experiment must confirm. The comparison should show how many additional cases are accepted by review and what human cost is no longer required.

Sonnet 4.5 reaches the operational threshold when it meets the team’s defined acceptance rate, regression limit, turnaround time, and cost. For localized fixes, the decision may favor it even if Opus earns a higher average score, when the difference does not reduce a relevant cost. In harder tasks, the result may be mixed: Sonnet for triage and bounded fixes, Opus for a defined queue of issues exceeding a complexity criterion.

Do not turn segmentation into a rule based solely on intuition. An initial policy may use signals such as the number of affected modules, the absence of a clear reproduction, or the need to analyze extensive traces. It must then be validated against the results. If those signals do not predict a sufficient improvement with Opus, they add complexity without improving the decision.

Provider announcements and technical cards can provide information about availability, configuration, and internal evaluations. They do not replace this test because the tools, repositories, prompts, budget, and success definition may not match those of your team.

08

Apply sensitivity analysis and state the limits

Repeat the comparison with different iteration budgets, context limits, and restricted tool permissions. A result that depends on a very high budget may not apply to an operation with strict limits. Likewise, an advantage observed with network access or a proprietary tool should not be attributed solely to the model.

Do not aggregate results from SWE-bench Verified, Terminal-Bench, or other evaluations with the internal result. It is valid to present them as methodological context if the difference in corpus and harness is identified, but not as comparable rows in the same table. Changing the tests, tool policy, or acceptance definition changes the task being measured.

This test also does not establish code security, authorization to deploy, continuous maintenance capability, or performance in every language. A patch accepted in an isolated environment may introduce risks not covered by the test suite. The production decision should retain review, continuous integration, and change-management controls.

Finally, document lost cases: excluded tickets, infrastructure errors, reviews without consensus, and unstable tests. Hiding them can make the conclusion appear more robust than it really is. Transparency is particularly important when the cost or acceptance difference between the two models is small.

09

A final template for making a repeatable decision

Before choosing, state the objective in writing: for example, reduce the cost per accepted maintenance fix without exceeding a limit on regressions or review time. Then run both models on the same frozen batch and publish enough parameters for the calculation to be repeated internally.

The conclusion should take a conditional form. For example: Sonnet 4.5 is the default option for the localized stratum because it reaches the acceptance threshold at a lower total cost; Opus 4.5 is reserved for the cross-module stratum if repetition confirms a sufficient reduction in rejections or human intervention. If the difference does not persist after repeating the batch, the responsible conclusion is that there is not enough evidence to pay a premium in that environment.

Review the decision when model identifiers, applicable prices, channel, available tools, or the composition of the issue queue changes. A reproducible evaluation is not a one-time purchase decision: it is a periodic control over a decision that depends on evolving systems and conditions.

Decision checklist for leaders

  1. 01Define the acceptable threshold for acceptance, regressions, and total cost.
  2. 02Build a representative, reproducible, stratified batch.
  3. 03Pin the models, channel, region, tools, context, and iterations.
  4. 04Run and review patches blindly using a predefined rubric.
  5. 05Calculate cost per complete success, not only cost per call.
  6. 06Repeat the experiment and adopt a routing rule only if the difference persists.

Open questions

  • Effective prices, billing concepts, and availability may vary by channel, region, contract, cache use, and run date; they should be checked at the start of every campaign.
  • Lifecycle documentation should be consulted again before using model identifiers, because availability and retirement dates can change.
  • No experimental results from Opus 4.5 and Sonnet 4.5 on the same corpus have been provided; this article therefore describes a protocol and does not claim an empirical advantage for either model.
  • Representativeness depends on the languages, repositories, issue classes, and tests included in the local corpus.
  • Blind human review reduces bias, but it does not eliminate disagreements or the possibility that the test suite misses relevant regressions.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction