The score does not belong to the model alone
A DeepSWE v1.1 score does not describe an isolated property of a language model. It describes the outcome of a complete experimental configuration: a model served by a provider, a reasoning-effort setting where that option exists, an agent harness, a tool set, context and time limits, instructions, a stopping policy, and an evaluation system. The observed unit is the patch committed by that configuration when it faces a particular task—not a textual model response or a general capability measured outside that environment.
This distinction is especially important when reading the DeepSWE leaderboard. The official site states that models run with mini-SWE-agent for consistency, but that choice does not remove every variable. The same model family may appear with different reasoning levels or providers; an update to the harness, prompt, tools, or infrastructure can also change the result without changing the underlying model. A direct comparison therefore requires checking that both rows use equivalent conditions and the same evaluation version.
The term pass@1 should not be read as a guarantee that an agent will solve the next issue in your own repository. In the DeepSWE methodology paper, pass@1 is defined as the per-task average of the success rate obtained across runs. It is an aggregation of observed performance on this collection and under this protocol. It is useful for summarizing experimental results, but it does not by itself identify the cause of each success, failure, or difference between configurations.
What DeepSWE is trying to measure
DeepSWE presents itself as a benchmark for software engineering agents working on original, long-horizon tasks. The authors’ documentation describes 113 tasks across 91 repositories. Each task combines repository context with a requested change and programmatic mechanisms for checking the expected behavior. The aim is not to retrieve an already published fix, but to produce a new solution for the task presented.
According to Datacurve’s presentation, reference solutions were written from scratch and were not copied or adapted from an existing pull request, commit, or public patch. Some tasks may be motivated by unresolved issues, but that motivation does not mean an upstream solution is available for use as a reference. This is a design claim from the creators; it matters because it seeks to reduce overlap with benchmarks built from historical issues, but it is not by itself proof that there is no contamination in training data or external tools.
The public format of the official repository is useful for auditing because it documents components such as metadata, prompts, Dockerfiles, tests or verifier logic, and reference solutions. This allows an evaluator to inspect a specific task rather than infer its difficulty from the repository name. Still, artifact availability does not automatically make every conclusion reproducible: repeating a leaderboard row also requires the relevant commit, the images or dependencies actually used, tool credentials and versions, and the pertinent execution logs.
What the benchmark observes—and what remains outside it
| Element | Signal it can provide | Conclusion it does not justify on its own |
|---|---|---|
| Committed patch | Ability to propose and apply a change under a closed harness | Maintainable quality in any production repository |
| Functional verifier | Compatibility of the change with behaviors encoded by the task | Complete coverage of requirements that were not encoded |
| Task success | Aggregate performance on the DeepSWE collection | Individual reliability on a future issue |
| Row comparison | Differences under documented, equivalent conditions | General superiority when protocol or version changes |
The patch, the clean container, and the verifier
For v1.1, Datacurve describes a procedure in which only the committed patch is evaluated in a separate, clean verification container. The official repository likewise documents extracting the commit as a patch and applying it in a pristine environment from that version onward. The design seeks to separate the final result from transient effects that may have occurred during the agent trajectory, such as uncommitted changes or manipulation of the test process inside its working environment.
The v1.1 documentation states that the CTRF report records each task-defining test by name. Under that approach, dropping tests or forcing an early exit should show up as missing or failed results rather than a pass. It also describes changes intended to correct dependency drift and unstable tests. These are reasonable improvements for evaluating a portable patch, but they should be read as benchmark implementation choices: every scoring system embeds assumptions about which files to restore, which commands to run, and which evidence counts as a result.
The separation between the agent container and the verification container has a practical consequence. A patch may work during one specific trajectory and fail when reapplied cleanly; in that case, a failure expresses a lack of reproducibility of the final change under the protocol. Conversely, a change that fulfills the task’s intent may fail if the restoration process, file substitution, or result detection interferes with the behavior being verified. Distinguishing those cases requires preserving enough artifacts to review the verdict.
The chain worth auditing for a task
- 01Identify the base repository commit, task identifier, and DeepSWE version.
- 02Record the agent configuration: model, provider, harness, tools, prompt, limits, and reasoning effort.
- 03Preserve the agent’s final commit and the patch extracted from it.
- 04Apply that patch to the clean environment defined by the task.
- 05Run the verifier and retain output, test report, logs, and exit code.
- 06Review which files were restored, replaced, or generated before attributing a failure to the agent.
What pass@1 and pass@4 score—and what must be disclosed
The DeepSWE paper defines pass@1 as the per-task mean success rate and pass@4 as the fraction of tasks solved in at least one of four attempts. The distinction matters: pass@1 reports the performance of an individual attempt under the run distribution used; pass@4 reflects the benefit of having several attempts available. They are not interchangeable measures, and rankings under one may order configurations differently from rankings under the other.
The authors document approximately four rollouts per task and 95% confidence intervals. Those intervals help prevent overreading small separations, especially when nearby configurations overlap. An interval, however, does not repair a methodological mismatch. If two results come from different verifiers, retry policies, execution deadlines, or exclusion criteria, their statistical uncertainty does not resolve the fact that the measured object has changed.
A responsible result sheet should disclose, at minimum, the DeepSWE version; model and provider; mini-SWE-agent version; reasoning level; number of rollouts; stopping policy; time and context limits; infrastructure date or version; and error handling. It should also separate the calculated score from excluded cases. Without that information, readers cannot tell whether a difference comes from problem-solving ability, service availability, budget decisions, or changes in the evaluator.
Counted failures, excluded errors, and selection bias
The methodology published by the authors states that timeouts count as failures. It also says that provider, network, or verifier errors may be excluded. This distinction makes operational sense: an external interruption is not necessarily equivalent to an inability to solve the problem. But exclusion creates a transparency obligation, because the classification of an event determines the denominator used in the published score.
A timeout may reveal a real limitation of the system intended for deployment, even if it does not reveal a semantic limitation of the model. A useful evaluation should therefore show both the primary rate under its rules and the number and nature of excluded runs, along with artifacts supporting each decision. Lumping a reproducible error in task setup under “infrastructure,” for example, would prevent third parties from assessing whether it is a benchmark defect or an exceptional environmental condition.
It is also worth asking whether an exclusion is decided before inspecting the patch and whether the rule is applied equally to every model. A retrospective or poorly documented policy can unintentionally favor one configuration. The available evidence here does not establish that this happens in DeepSWE; it is an audit condition that should be checked before using the score for procurement, vendor selection, or deployment decisions.
Treatment that should be made explicit
| Event | Under the declared methodology | Minimum evidence needed for review |
|---|---|---|
| Timeout | Counts as a failure | Configured limit, logs, and observed duration |
| Provider error | May be excluded | Service response, timestamp, and applied criterion |
| Network error | May be excluded | Connectivity record and controlled rerun |
| Verifier error | May be excluded | Reproducible failure, verifier version, and technical explanation |
| Failed test | Task failure unless a reclassification is justified | Full output, patch, and container state |
Why v1 and v1.1 should not be combined without rerunning
Datacurve states that v1.1 retains the same tasks but changes execution and evaluation. The described changes include verifying the committed patch in a clean container, CTRF reporting, and fixes relating to dependency drift and unstable tests. Even if the set of task statements remains unchanged, a modification to the scoring environment can change which patches pass and which fail.
That prevents treating a v1 score and a v1.1 score as points in one continuous series without a clear warning. It would not be valid to attribute the whole difference to a model improvement or regression if the method of reconstructing the environment or checking tests also changed. The strongest way to compare versions is to rerun the same configuration under each protocol and publish task-level results alongside changes in verdicts.
The same caution applies to comparisons within v1.1 if the leaderboard evolves. A public table is a valuable record of declared runs, not sufficient proof of historical equivalence. Anyone making a high-impact decision should freeze a benchmark and harness commit, retain the container images, and run a controlled configuration matrix.
The dispute over test modifications
The independent review by Epoch AI raises a specific limitation of DeepSWE v1.1. According to that review, its authors manually confirmed at least 23 false negatives among the 113 tasks. They explain that 18 of those cases are related to agents modifying existing tests and to subsequent verifier setup that restores or replaces files. The stated result is that an agent may have produced a functionally correct change for the task’s intent and nevertheless received a failing verdict because of the interaction between its patch and the evaluation protocol.
This finding should be attributed to Epoch AI and read within its stated scope. The review says it stopped after exceeding its threshold for considering the benchmark flawed; it therefore does not provide a definitive estimate of the total false-negative rate or establish that every DeepSWE result is invalid. Nor does that figure alone allow an inference about how much leaderboard positions would change: doing so would require rerunning affected cases under a corrected version and publishing verdict reassignment by configuration.
There is also a real tension between two objectives. The evaluator wants to prevent an agent from obtaining a pass by changing the task’s tests. But a rule that restores or replaces files must distinguish manipulation that evades checking from a test edit that is part of a legitimate change or that interacts with the verifier in an unintended way. The official v1.1 documentation maintains that the test report detects missing results and early exits. The audit suggests that this safeguard does not prevent all identified false negatives. With the available sources, both claims describe different levels: the intended design and a set of behaviors observed during review.
Minimum protocol for a suspected false negative
- 01Pin the task commit, container, and exact v1.1 version under examination.
- 02Reproduce the original verdict using the committed patch, with no manual changes.
- 03Inspect the diff to separate product changes, test changes, and generated files.
- 04Record verifier operations that restore, replace, or ignore files.
- 05Run checks that evaluate the requested behavior without creating an alternative path to a pass.
- 06Classify the case with a public rationale: patch failure, false negative, ambiguous behavior, or environment defect.
What a high score does allow you to conclude
A high DeepSWE v1.1 score is evidence that a particular configuration was able to produce patches accepted by the protocol on a collection of original software engineering tasks. When the result is accompanied by artifacts and reproducible conditions, it provides a useful signal for preselecting agents intended for autonomous repair and development tasks involving repositories, commands, and tests.
The signal is more informative than an evaluation limited to isolated code snippets because it includes repository navigation, multi-file modifications, tool use, and functional verification. It can also help expose practical differences between configurations that appear similar on more saturated benchmarks. Its value is conditional, however: it depends on whether the task population, harness constraints, and verifier resemble the use case on which a decision will be based.
The prudent conclusion is not that the agent “can maintain production software.” It is that it has shown measured ability to solve a fraction of this benchmark’s tasks under a defined protocol. That formulation preserves useful information while avoiding the transfer of an experimental result to domains DeepSWE does not directly observe.
What it does not allow you to conclude—and how to read a leaderboard row
The benchmark does not by itself establish reliability in your own repository, change safety, design-review quality, compatibility with internal policies, or the ability to debug systems connected to production. Nor does it fully measure operational cost: a comparable score may require different token counts, wall-clock time, retries, or human supervision. Enterprise autonomy additionally requires permissions, isolation, dependency controls, review, observability, and rollback procedures.
A row also does not prove that the agent did not find a solution exploiting a gap in the verifier, nor that every failure represents an inability to solve the requirement. Epoch AI’s review makes this latter caveat particularly relevant in the cases it examined. Conversely, a partial audit of false negatives does not establish that the benchmark has no value. The correct interpretation is that measurement error and disagreement between intent and verdict must be included in the risk analysis.
To compare Grok 4.6, DeepSeek V4.1 Flash, or other systems on DeepSWE, a technical buyer should first filter rows to the same version, the same harness, and equivalent inference conditions. They should then inspect intervals, excluded errors, and artifacts from boundary cases. Finally, they should complement preselection with an internal evaluation on representative repositories, using security policies and costs close to the intended deployment. DeepSWE can be an input to that decision, not its replacement.
Checklist for reading a published row
| Question | Why it matters |
|---|---|
| Which DeepSWE and verifier version was used? | Avoids combining results obtained with different instruments. |
| Which model, provider, and reasoning effort are listed? | Defines the configuration that was actually measured. |
| Which harness, prompt, tools, and limits were used? | Supports more careful attribution of the outcome to the complete system. |
| How many rollouts were run, and which metric is shown? | Separates single-attempt performance from the benefit of retries. |
| How many cases were excluded, and why? | Allows assessment of the denominator and operational availability. |
| Are task-level logs, patches, and verdicts available? | Makes it possible to investigate successes, failures, and ambiguous cases. |
Open questions
- The provided sources do not establish whether the 23 cases identified by Epoch AI reproduce across all later v1.1 revisions or what their full effect would be on each leaderboard position.
- No subsequent technical response from Datacurve is provided that resolves or disputes the false negatives described by Epoch AI on a case-by-case basis.
- The sources do not support inferring the false-positive rate: patches that are approved without fulfilling a task’s intent.
- Exact equivalence between specific leaderboard rows depends on configuration details and execution artifacts that must be checked in each publication.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction