Ilustración editorial para SWE-Bench Verified: qué mide realmente un resultado y por qué no basta para elegir un agente de código
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What Question SWE-Bench Verified Answers—and What It Does Not

SWE-Bench is an evaluation based on historical GitHub issues and their associated corrective changes. Its Verified subset contains 500 human-reviewed instances intended to improve evaluation reliability. In practical terms, it asks a bounded question: given the supplied context for an issue in a specific repository, can a system propose a patch that passes the tests defined for that task in a reproduced environment?

The answer is valuable because it connects code generation to an executable outcome, rather than only to human preference or textual similarity to an expected solution. An agent must navigate a codebase, interpret a description, edit files, and generate a modification compatible with the environment’s tests. However, the metric reduces that process to a proportion of tasks meeting a binary resolution criterion.

On its own, it does not answer whether an agent can operate with responsible autonomy in a real repository. It does not establish that the agent can prioritize an issue queue, ask for clarification on ambiguous requirements, decide that code should not be changed, review a third-party contribution, manage secrets, coordinate a deployment, respond to an operational alert, or accept responsibility for a regression. Those activities depend on people, policies, systems, and contexts that are not fully represented by a closed instance.

For that reason, a score should be read as evidence about a capability evaluated under a particular protocol, not as a general ranking of engineering tools. In the Benchmarks, Compare, and Discover sections, it can reasonably serve as one signal among several, provided its conditions of production are retained and an isolated number is not turned into a promise of operational outcomes.

02

Anatomy of a Task: From a Historical Issue to an Evaluated Patch

The original construction of SWE-Bench starts with resolved GitHub issues from 12 open-source Python repositories and links them to the corresponding code change. At a conceptual minimum, an instance brings together the issue, the historical repository state on which work will be performed, and the project’s reference change. The benchmark turns that historical material into a task that a system can attempt to resolve through a code edit.

It is important to distinguish the reference change from the win condition. The goal is not necessarily to reproduce the original human patch character for character. An alternative patch may be accepted if, when applied to the instance environment, it satisfies the intended correctness tests and does not break preservation tests. This distinction matters because it prevents treating the benchmark as an exact-answer retrieval exercise.

Verified adds a human filtering layer over SWE-Bench instances. Project documentation describes the set as a human-validated subset of 500 instances. The selection aims to remove cases whose evaluation is not sufficiently reliable; nevertheless, the subset does not turn every problem into an exhaustive representation of software engineering, nor does it eliminate every possible interpretive ambiguity.

A reproducible run requires more than the issue text. It should specify the data revision, the identifier of every instance, the environment image or definition, the harness version, the produced patch, and execution limits. The consulted data distribution is served from a branch that can change. Therefore, citing only the branch name does not pin the same set for future repetitions. A complete revision should be fixed, and the access date should be recorded.

03

How Resolution Is Defined, and Why the Final Percentage Hides Methodological Choices

According to the evaluation description, the agent does not see the tests. The harness evaluates its patch through two groups: FAIL_TO_PASS tests and PASS_TO_PASS tests. The first group represents behavior the change must fix; the second checks that the edit has not unintentionally broken unrelated parts of the codebase. For an edit to count as a complete resolution, both groups must pass.

This definition is more demanding than checking that code compiles or that a single new test passes. It also makes it possible to compare functionally different patches without requiring equality with the historical change. Yet a final percentage still hides choices: how many instances entered the denominator, what happened to infrastructure errors, how much time each task received, how many samples were generated, and what policy selected one among them.

The official harness documents parameters for choosing the dataset and instances, setting a timeout, and separating runs. It also applies patches, executes tests, and calculates results. Consequently, reporting only a resolution rate without publishing the command or an equivalent configuration makes it difficult to verify whether two figures used the same tasks, limits, and evaluation mechanism.

The opposite oversimplified conclusion should also be avoided: that a binary metric has no value because it does not cover the entire development lifecycle. The metric can provide concrete evidence about issue repair under defined conditions. The relevant task is to delimit its scope and require enough traceability for someone else to inspect the conditions behind the claim.

Choices that the same resolution rate can conceal

FieldAudit questionEffect on interpretation
DenominatorWere all 500 instances evaluated, or only a subset?A rate over filtered tasks does not necessarily represent the complete set.
SamplesWas there one patch or several attempts per task?More attempts can increase the likelihood of finding a valid patch.
SelectionHow was the evaluated patch selected from multiple outputs?The selection rule may use different information or incur different cost.
Time and computeWhat limits on steps, calls, and time were applied?The figure alone does not express the system’s efficiency.
Execution failuresHow were timeouts and infrastructure incidents handled?Excluding or retrying them changes the effective denominator.
04

The Seven Minimum Fields for Comparing Two Published Results

A responsible table does not need to disclose every internal detail of a system, but it should make one run distinguishable from another. The first field is the exact identity of the set: name, variant, pinned revision, and the instance list or selection rule. “SWE-Bench Verified” is not enough if the result was produced on a sample, a modified copy, or a different data version.

The second field is the model: provider, name, version or identifiable date where available, and inference parameters that affect the outcome. The third is the agent scaffolding: framework, version, and planning or editing strategy. The same model can produce different results if the tool loop, context formatting, error handling, or stopping criterion changes.

The fourth field describes enabled tools: terminal, local search, file editing, test execution, network access, and any external retrieval. The fifth declares the budget: limits on steps, model calls, tokens when available, time per instance, and compute resources. The sixth states the protocol: number of samples, temperature or another generation setting, retries, and the selection rule. The seventh provides the artifacts that enable verification: configuration, sufficient logs, patches or predictions, and harness output when they can be shared.

The project page distinguishes a general leaderboard that combines heterogeneous systems from a model comparison using a common configuration based on mini-SWE-agent. This separation is a methodological warning: a ranking of full systems does not isolate the model’s effect, while a common configuration can help with that specific comparison. The project also warns that versions 1.x and 2.x of mini-SWE-agent are not necessarily comparable.

Not all of these details will be available in a release note. In that case, the result is not automatically disproven, but it should be labeled incompletely specified. The rigorous response is not to fill gaps with assumptions about the provider or rank heterogeneous figures as though they came from a controlled experiment.

Minimum record for a published figure

FieldWhat should be statedWarning sign
DatasetVariant, revision, and included tasksOnly the benchmark name is stated.
ModelIdentity and version or dateA commercial name without an identifiable version.
AgentFramework and versionThe scaffolding that uses the model is omitted.
ToolsAvailable capabilities and constraintsIt is unclear whether network, terminal, or tests were available.
BudgetPer-task limits and cost or resources when knownQuality is compared without equivalent limits.
ProtocolSamples, retries, and selectionThe method for choosing the final output is not explained.
EvidenceConfiguration and verifiable artifactsOnly an aggregate claim is available.
05

What Can Inflate or Limit a Score

Task selection is the first factor. A subset chosen by difficulty, environment availability, or previous successes does not necessarily preserve the distribution of the full set. It also matters whether instances that time out, fail to build an image, or encounter infrastructure problems are excluded. A report should, as far as possible, separate an agent failure from an environment failure and explain how both affect the aggregate result.

Retries and multiple samples deserve specific attention. Trying several patches for each issue can be a legitimate technical decision, especially if it reflects the system’s intended use, but it changes the practical unit of evaluation: success is no longer measured for a single attempt. The maximum number of attempts, whether the agent is restarted, and whether failed runs are launched again should all be reported.

The information accessible to the system changes the nature of the task. Hidden tests reduce one direct route for adapting to the expected answer, but they do not eliminate other differences: depending on the protocol, the agent may have search, execution tools, local documentation, or external access. A valid comparison requires knowing which resources were enabled and whether they were equal for all systems being compared.

There is also a temporal uncertainty. OpenAI has stated its view that SWE-Bench Verified no longer measures frontier programming capabilities and has identified the risk of contamination arising from public availability of historical problems and solutions. This is OpenAI’s assessment and recommendation, not an independent measurement that can quantify contamination for every model. Even so, it requires caution when interpreting a recent improvement as general progress without examining potential exposure to the data.

Epoch AI’s analysis raises another limitation: the benchmark is concentrated in familiar repositories and relatively bounded fixes. That is a secondary interpretation, not a property that should be presented as a definitive fact about every instance. It nevertheless supports a useful question: does the organization’s maintenance portfolio materially resemble these historical tasks from Python repositories? If not, the expected transfer of the signal will be limited and uncertain.

06

Why a Resolved Issue Does Not Demonstrate Autonomous Maintenance

In a real repository, resolving an issue begins before writing a patch. Teams need to triage, reproduce the problem, estimate impact, identify dependencies, negotiate requirements, and decide priorities. A benchmark task supplies a historical formulation and a prepared test criterion; in normal operations, those inputs may be absent, contradictory, or changing during the investigation.

After the patch, other activities also matter and are not sufficiently covered by a resolution rate: peer review, security analysis, licensing, backward compatibility, migrations, performance, observability, change approval, and deployment. The benchmark’s preservation tests are an important safeguard within an instance, but they are not equivalent to all of an organization’s validation processes or to the effects of integrating a change with current branches, services, and users.

Operational accountability is another boundary. An agent may generate a modification that passes harness tests and still require human supervision to decide whether it should merge, when it should deploy, and how it should be reverted. Therefore, purchasing, adopting, or granting write access to a tool should not depend solely on a SWE-Bench Verified percentage. It should incorporate access controls, review, traceability, isolation, and tests specific to the organization’s own environment.

This does not make the benchmark irrelevant to engineering leaders. It can help select hypotheses for a subsequent test: for example, if a system shows the ability to edit and validate patches on historical tasks, it may deserve a controlled evaluation on low-risk internal issues. The correct transition is from benchmark evidence to a local experiment, not from a benchmark to production autonomy.

Ten-minute audit protocol

  1. 01Identify whether the figure covers the complete Verified set, a subset, or a variant; record the declared data revision.
  2. 02Check the model identity, its date or version, and the agent framework version.
  3. 03Find which tools the agent had, especially test execution, terminal access, network access, and external retrieval.
  4. 04Record limits on time, steps, calls, and the number of samples per instance.
  5. 05Determine the retry policy and how the final patch was selected.
  6. 06Verify that the success criterion includes applicable correctness and preservation tests.
  7. 07Examine how timeouts, image errors, and infrastructure failures were handled.
  8. 08Distinguish a self-run evaluation with artifacts from a claim without reproducible evidence.
  9. 09Avoid directly comparing mini-SWE-agent configurations that the project warns are not necessarily comparable.
  10. 10Conclude with a label: comparable, partially comparable, or not comparable; do not force a numerical ranking when essential fields are missing.
07

How to Transfer the Signal to a Brief Test in Your Own Repository

It is not necessary to reproduce all of SWE-Bench to obtain information that is closer to local reality. A brief test can use a small set of already closed issues or deliberately prepared changes, provided that responsible teams define inclusion criteria, permitted access, and the evaluation method in advance. The purpose is not to create a new public leaderboard, but to reduce uncertainty around a specific technical decision.

The design should separate development tasks from evaluation tasks. The person or team preparing cases may retain acceptance tests that are not visible to the agent, where feasible and appropriate. Every case should include an isolated environment, a pinned repository state, and explicit limits on time, cost, and tools. The system should not receive access to production credentials or be allowed to make changes outside the controlled environment.

Measure more than one dimension. In addition to whether tests pass, record time to patch, number of human interventions, quality of the explanation, compliance with repository conventions, review findings, and security or process incidents. A small sample does not support broad inferences; it can, however, reveal clear incompatibilities, unexpected costs, or classes of tasks where the system needs too much supervision.

The most useful comparison holds the protocol constant. If two systems are tested, aim to give them the same cases, time window, tool access, and limits. If the agent changes or one system receives a larger budget, report that change as part of the result rather than attributing the entire difference to the model.

A bounded and safe local test

  1. 01Select a small number of representative cases and classify them by type and risk.
  2. 02Pin commits, dependencies, and isolated environments before running agents.
  3. 03Define acceptance tests and an independent human review of each patch.
  4. 04Set minimum permissions: no secrets, no production, and no writes outside the test environment.
  5. 05Run each system with documented budgets and tools.
  6. 06Record results, costs, timing, environment failures, and reasons for rejection.
  7. 07Make a decision based on observed patterns and operational limits, not on one aggregate rate.
08

Final Record: What Can Be Claimed Rigorously

A strong claim takes a bounded form: “On the declared revision of SWE-Bench Verified, with this model, this agent version, these tools, this budget, and this protocol, the run obtained this rate of instances that passed the harness criterion.” If artifacts are available, it can add that the result is auditable or reproducible under the published conditions. If they are missing, the appropriate statement is that the claim cannot be fully verified from the available information.

It is not rigorous to turn that statement into “the model resolves that percentage of real bugs,” “it is the best coding agent,” or “it can maintain a repository without supervision.” Those conclusions expand the population, context, and responsibilities without equivalent evidence. Even an impeccable benchmark run only answers for the tasks and protocol that were actually evaluated.

Primary documentation provides clear foundations for this reading: Verified is a human-validated subset of 500 instances; resolution requires passing both correctness and preservation tests; and harness configuration is a material part of the run. At the same time, uncertainties remain: public availability of tasks may affect temporal validity for certain models, agent configurations change, and a historical task does not reproduce all social and operational mechanisms of maintenance.

The practical decision is to retain both ideas. SWE-Bench Verified can be a useful technical signal and more concrete than an anecdotal demonstration. It is not a guarantee of autonomy, security, net productivity, or suitability for a particular repository. Anyone publishing, comparing, or buying on the basis of these results should make visible both the conditions that turn a number into evidence and the uncertainties that prevent it from becoming a promise.

Recommended language for communicating a result

SituationRigorous wordingWording to avoid
Documented runIt obtained a resolution rate under the stated protocol and budget.It resolves real issues at that percentage.
Comparison under equal conditionsIt outperformed another system in this common configuration.The model is generally superior.
Result without complete configurationA figure has been reported, but information needed for direct comparison is missing.The figure proves the model’s performance.
Internal useIt justifies a controlled trial on local tasks.It justifies production autonomy.

Open questions

  • The data distribution branch is mutable; reproducibility requires a pinned complete revision and an access date, not only the branch name.
  • Potential contamination from public data is a concern expressed by OpenAI; the supplied sources do not allow its effect on a specific model or result to be quantified.
  • The transfer of results to an organization’s particular repositories, languages, processes, and risks cannot be directly inferred from a SWE-Bench Verified rate.
  • Without configuration, logs, and artifacts for a published run, it is not possible to determine whether a score difference comes from the model, agent, budget, or protocol.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction