The misleading number: price per resolved issue does not explain itself
Presenting an agent’s result as “cost per resolved issue” appears to turn a technical evaluation into a simple economic decision. Yet that figure is a ratio between variables that may have been defined in very different ways. The numerator may contain only the tokens from one final call, or all calls started during a trajectory. The denominator may be the number of patches that pass the verifier, the number of issues initially selected, or a result chosen from several attempts. Without those definitions, two identical figures do not necessarily describe comparable operations.
This distinction matters especially in SWE-bench. The task begins with real issues associated with repositories and requires producing a change that can be verified in an environment prepared for that purpose. The observable cost is therefore not limited to writing a patch. A system may inspect files, request additional context, run tools, restart a strategy, or exhaust a limit without producing a valid solution. All of those events may consume resources even when the instance is not counted as resolved.
The right economic question depends on the decision. To budget a full run, the relevant measure is the average cost per attempted issue. To assess the performance of a workflow that delivers only verified changes, total cost divided by strict successes may be more useful. To decide whether a more expensive configuration is worthwhile, the relevant comparison is usually the additional cost per additional success relative to a baseline. And if several trajectories are authorized and one is retained, the cost of the full policy must be measured, not only the cost of the selected trajectory.
This does not mean that a summary metric is useless. It means it should be read as the cover page of a cost scorecard. That scorecard must identify the dataset version, the subset actually run, exclusions, the number of attempts, the selection rule, the stopping protocol, consumption categories, and which part of the evaluation infrastructure was included. Without those elements, the number may be a valid internal observation, but it is not a sufficient basis for comparing agents or estimating an automation deployment.
What is paid for during a run
The first component is inference. In an API, the usage record should separate input and output tokens when the provider bills them differently. If the platform reports cached input or billable reasoning tokens, those categories should also appear separately. OpenAI documentation, applicable only to systems that use that API, distinguishes these categories and notes that requesting multiple completions consumes additional tokens. Neither that terminology nor its pricing structure should be automatically extrapolated to other providers.
The second component is auxiliary calls and tools. An agent may search, summarize files, generate tests, review diffs, or request new completions after running commands. Some tools have no direct API price, but they increase the context of later calls or use their own compute. If an external tool charges per use, it should appear as a separate line item. If no monetary cost is assigned to it, its invocation count and resource consumption should at least be recorded so that another organization can value them using its own rates.
The third component is evaluation. The SWE-bench harness prepares Docker images, applies patches, and runs tests with per-instance time limits. Its documentation also covers CPU, memory, storage, and caching requirements. Building or retrieving images, running containers, and retaining artifacts can represent a material cost, especially in large campaigns, even though they are not inference costs. Combining them without a breakdown makes it impossible to know whether an improvement comes from the agent or the infrastructure; omitting them from an operating budget can understate the actual expense.
It is also useful to separate amortized cost from marginal cost. Preparing a shared image for many instances does not cost the same per run as rebuilding it from scratch. Similarly, a result cache can avoid later work. These efficiencies are legitimate when documented, but they should not be presented as though every issue required the same marginal expense. A strong scorecard provides both views: the cost of the observed campaign and the rules used to allocate shared costs.
Minimum cost items worth separating
| Item | What to record | Risk if omitted |
|---|---|---|
| Inference | Input, output, cached, and billable reasoning tokens; model, endpoint, and pricing date | A cost driven by context, retries, or undisclosed rates is attributed to a model |
| Tools | Calls, paid services, execution time, and artifacts | Auxiliary work needed to produce the patch is made invisible |
| Evaluation | Images, containers, CPU, memory, storage, tests, and waiting times | The operating budget does not cover the cost of verifying changes |
| Shared costs | Method for amortizing images, caches, and preparation | A comparison combines campaigns with unequal reuse |
Four denominators for four different questions
The first metric is average cost per attempted issue: total campaign cost divided by every issue for which the protocol was started. It is the most suitable measure for estimating the expense of processing a similar work queue because it includes successes, failures, invalid trajectories, and cases exhausted by limits. It should be clear whether an issue excluded before the agent starts belongs to the universe or not. An exclusion made after execution should not erase its cost.
The second metric is cost per strict success: total cost of trajectories included under the policy divided by the number of issues whose patch passes the defined verification process. It describes how much spending was needed, on average, to obtain a validated result under that protocol. It can rise even if cost per attempt falls when the resolution rate declines, which is why it should not be published without both underlying quantities.
The third metric is the incremental cost of increasing resolution. If configuration B resolves more issues than configuration A, it is calculated as the difference in total cost between B and A divided by the difference in successes. This ratio answers a marginal decision: how much each additional success costs under a more ambitious configuration. It is interpretable only if both systems run on the same set, with equivalent success criteria and evaluation resources.
The fourth applies to policies with several attempts per issue. If pass@k, retries, or later selection are allowed, cost must include the trajectories generated for each issue, including those not selected. Reporting the cost of the best trajectory found amounts to reporting a conditional result that does not reproduce the expense of finding it. The publication should clarify whether one trajectory was run per instance, several independent trajectories, or an adaptive decision tree.
The protocol changes both the bill and the meaning of the result
A limit on steps, time, calls, or monetary budget is part of the intervention being evaluated. Increasing it may allow some difficult issues to receive more exploration, but it can also concentrate a substantial share of spending in a long tail of unresolved cases. For that reason, in addition to the mean, it is useful to publish the median, percentiles, and per-instance distribution. A low mean can coexist with a small number of exceptionally expensive cases that would be unacceptable in production.
The stopping policy must be explicit. A run may stop when it obtains a patch, fails to find relevant files, exceeds a budget, fails tests, or reaches a number of iterations. Each option changes both cost and probability of success. Retries require the same precision: it should be stated what triggers them, whether they inherit context, whether they reuse cache, and whether discarded attempts are counted. Calling a sequence of internal restarts “one attempt” can conceal a material difference in consumption.
Pass@1 and pass@k answer different questions. Pass@1 approximates the performance of a single opportunity under a fixed configuration. Pass@k describes the possibility that at least one of several opportunities produces a success, but it is not equivalent to the cost of one opportunity. When later selection is used, the signal used to select should be documented, along with whether selection consumed additional model or compute resources and whether it had access to test results. Selection informed by the verifier can be useful for research, but it should not be confused with a policy available before verification.
The most informative comparison usually fixes common constraints: the same set of instances, the same harness, the same evaluation limits, and, when the goal is economic, a comparable maximum budget per issue. Cost-versus-resolution curves can then be shown. This presentation reveals whether the resolution gain appears at a gradual cost or depends on a minority of very expensive trajectories. It also avoids attributing to agent quality what may be due to allowing it to spend more.
Process for turning a summary figure into an auditable scorecard
- 01Fix the dataset revision, the list of eligible instances, and every exclusion together with its reason.
- 02Record, for each instance, every trajectory started, its stopping condition, retries, and final state.
- 03Aggregate inference consumption by category and apply the pricing table in force on the declared date.
- 04Measure tool use and evaluation infrastructure separately, including shared resources and the allocation criterion used.
- 05Calculate metrics per attempted issue, per strict success, marginally, and per pass@k policy where applicable.
- 06Keep results, aggregate logs, and enough configuration information for a third party to reproduce totals without exposing secrets.
Inference and evaluation: separating them does not mean ignoring either
Separating inference from evaluation makes it possible to answer two questions that a single figure blends together. The first is how much it costs the agent to propose a change. The second is how much it costs to determine whether that change passes the benchmark protocol. In SWE-bench, evaluation requires applying the prediction and running tests in a repository environment. The harness documents image preparation, container execution, time limits, and caching options; therefore, treating verification as a free operation would be methodologically incomplete.
Separation does not require choosing a single convention. For model research, it may be reasonable to report inference cost first and, alongside it, the evaluation cost of the campaign. For planning an automated repair service, total cost of ownership is more relevant: inference, orchestration, tools, test compute, storage, and human review when it is part of the workflow. The key is not to add some cost items for one agent while omitting them for another.
There is a practical uncertainty: infrastructure costs depend on region, provider, reserved capacity, concurrency, and retention policy. The available sources describe harness components, but they do not establish a universal rate for running them. A rigorous publication should therefore provide physical units, such as container time and allocated resources, in addition to any local monetary conversion. Another organization can then recalculate the amount under its own contracts.
Experiment artifacts are essential to this separation. The SWE-bench experiments repository covers predictions, execution logs, traces, and per-instance results. Sharing or summarizing these artifacts in a consistent structure makes it possible to check which patches were evaluated, detect instances without a result, and reconcile cost totals with trajectories. Auditability does not require revealing credentials, confidential prompts, or protected data; it does require that exclusions and aggregations do not prevent the accounting from being reviewed.
The minimum publication scorecard and the operational decision
The scorecard should start with the experiment identity: dataset variant and revision, number of eligible, attempted, and excluded instances, together with the reasons for exclusion. This matters because SWE-bench offers several variants, and project documentation identifies SWE-bench Verified as a set of 500 instances. Merely naming “SWE-bench” is not enough to know which population was evaluated. The harness identifier, relevant images or configuration, and verification rules should also be retained.
Next, the model, provider or endpoint, region when it changes the price, price lookup date, and currency should be recorded. The accounting should show input, output, cached, and reasoning tokens when those categories exist for the provider used, as well as auxiliary calls. It should include the per-issue limit, stopping policy, retries, and selection method. Means should be accompanied by the per-instance distribution and counts of failed, exhausted, invalid, or unverifiable trajectories.
The scorecard ends with two totals: inference and evaluation. For each, it should indicate which items it includes, what is excluded, and how shared costs are allocated. If a per-success figure is published, the numerator total should reconcile with the preceding line items and the denominator with the observed strict successes. If there are multiple samples per instance, the reported cost must be the cost of generating and choosing among all of them, not the cost of the winning sample.
When exploring models at an early stage, cost per attempted issue and a resolution curve under fixed budgets are usually the most useful metrics. To optimize an agent, add incremental cost per additional success and the distribution of expensive cases. To budget automation, the reference is total cost of ownership per incoming issue, including verification and the human intervention actually required by the process. None of these measures replaces the others: each addresses a different risk.
The cautious conclusion is that an agent does not necessarily reduce the cost of resolving issues simply by obtaining a higher resolution rate or displaying a low amount associated with its successes. It may shift spending toward more trajectories, more context, more test compute, or later selection. A defensible comparison declares that shift. With a complete scorecard, a team can decide whether to pay for more successes, limit exposure to expensive cases, or adopt a configuration that offers more predictable economics.
Which metric to use for each decision
| Decision | Primary metric | Information that must not be missing |
|---|---|---|
| Explore configurations | Cost per attempted issue and resolution under a fixed budget | Limits, failures, spending distribution, and an identical dataset |
| Improve an existing agent | Incremental cost per additional success | Baseline, success differences, selection policy, and retries |
| Budget operations | Total cost of ownership per incoming issue | Inference, evaluation, tools, infrastructure, and human review |
| Compare published results | Cost per strict success alongside cost per attempt | Denominator, pass@1 or pass@k, exclusions, prices, and date |
Open questions
- The supplied sources describe the dataset and harness components, but they do not provide a universal rate for CPU, storage, containers, or auxiliary services; those cost items depend on each organization’s environment.
- The categorization of input, output, cached, and reasoning tokens is supported specifically by OpenAI documentation and should not be generalized without checking the documentation of the specific provider.
- The availability and detail of traces, invoices, or artifacts may be constrained by secrets, licenses, internal data, or retention policies; an audit may require verifiable aggregates rather than raw data.
- A benchmark evaluation does not by itself determine cost or success rate for production issues, where repositories, tools, security requirements, and human review differ.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction