A reproduction test, not an isolated programming task
RECLAIM is a benchmark designed to measure whether an artificial intelligence agent can reproduce a specific result described in a machine learning paper. The preprint presents a set of 100 NeurIPS 2025 papers and defines one task for each: using the paper and any materials released by its authors, the agent must work within a GPU-hour budget set in advance.
The scope of the work distinguishes this from a conventional programming test. To reach a result, an agent may need to install software, resolve errors, understand the method, run experiments, and check its results. The evaluation aims to cover that chain of tasks, rather than simply determine whether the system produces code that looks correct.
The benchmark also specifies in advance which result must be reproduced and what conditions count as success. That is an important choice: without a predefined goal, comparisons between agents could depend on differing judgments about what constitutes a satisfactory reproduction. The preprint’s abstract does not detail the exact rules used for each paper, so it is not enough to reconstruct how the criterion was applied in every case.
Three tiers based on what the authors released
RECLAIM groups tasks according to the availability of resources. In the Run tier, the code, data, and model weights are available. In Retrain, the weights are missing, so the agent must train the model. In Reimplement, the code is unavailable and the agent must write an implementation. According to the preprint, the materials released by the authors determine the difficulty tier.
This classification helps put the results in context: tasks do not all start from the same point. Running an already prepared system, rebuilding its weights through training, and implementing a method from scratch require different kinds of work. A single overall success rate, without a breakdown by tier, would therefore hide important differences.
The classification does not mean that every task within a tier has identical requirements. Papers may describe different methods and experiments. The available abstract does not list the resources for each paper or explain how much their requirements vary. The tiers therefore describe the availability of certain materials, not a perfect equivalence in difficulty.
How to read the RECLAIM tiers
The table summarizes the three tier definitions given in the preprint. It is not an independent ranking of the difficulty of individual papers.
| Tier | Resources described | Work the agent must take on |
|---|---|---|
| Run | Code, data, and weights | Run the released materials and obtain the specified result |
| Retrain | Weights are missing | Train the model, in addition to completing the rest of the task |
| Reimplement | Code is missing | Write an implementation of the method in an attempt to reproduce the result |
The results show a gap between tiers
The preprint reports that four agents were tested once on each paper. The best agent in each tier reproduced 41% of Run papers, 27% of Retrain papers, and 15% of Reimplement papers. In this set of tasks, and under the described procedure, success declined as resources needed to rebuild the system were removed.
These figures are benchmark results, not a universal measure of the capabilities of any agent. Nor should they be read as a direct comparison between agents without further information: the abstract reports the best result in each tier, but does not give the four systems’ names, their individual scores, or enough detail to reconstruct the variability across attempts.
The paper also reports that failed attempts used, on average, 29% of the assigned budget. The preprint’s interpretation is that many agents stopped with budget still available. This suggests that the compute limit alone does not explain every failure; by itself, however, the figure does not establish why each attempt stopped or what change would have been enough to achieve a reproduction.
Another frequent error was implementing the method without checking any part of it against the values reported in the paper. The abstract records this error in 63 of 400 runs. This observation points to a distinction between producing a plausible implementation and checking it against evidence: writing out the method does not guarantee that the agent reproduced the relevant conditions.
A cautious way to read the success rates
- 01Identify the resource tier: Run, Retrain, or Reimplement.
- 02Read the rate as the best agent’s result in that tier, not as the average across all agents.
- 03Keep in mind that the report describes one run per agent and paper.
- 04Do not turn the percentage of reproduced results into a claim about the overall validity of the papers.
What it evaluates—and what it leaves out
RECLAIM evaluates whether an agent can reach a previously selected result using the available materials and within a computational budget. According to the abstract, a separate instance of a language model grades the runs based on logs and outputs, rather than relying on the agents’ written reports. The goal is to assess what happened during execution, not merely what the system says it did.
That also defines the limits of the conclusion. Reproducing a specific result does not automatically verify every methodological choice in a paper, the quality of its data, the robustness of its analyses, or the validity of its scientific conclusions. Conversely, failure to complete a benchmark task does not by itself show that the original result is wrong: there may be technical, implementation, or resource-related causes that the abstract does not break down.
This distinction matters for readers and teams that want to use the benchmark as a signal of progress. A low success rate may reveal that agents struggle to reconstruct experiments from incomplete information; without additional evidence, it should not be treated as a verdict on the paper being evaluated.
Reproducing the evaluation also requires details
To compare agents independently, knowing the benchmark’s size and success rates is not enough. Researchers would need access to the paper selection, the target result for each task, the operational success criterion, the specific GPU budgets, and the time limits. The agents and their configurations, the instructions they received, the environments, the available materials, and the logs or outputs used to grade each run also matter.
The preprint abstract confirms that RECLAIM sets the target result, success criterion, and GPU-hour budget for each paper, and that evaluation is based on logs and outputs. However, the information provided here does not specify the budget values, how the 100 papers were selected, the names of the four agents, or whether the environments and scripts needed to repeat the evaluation have been published. Those points should be checked in the paper and its accompanying materials before attempting a more detailed comparison.
The presentation describes RECLAIM as a benchmark that can be rebuilt each year from new conferences. This offers a possible way to track changes in agent capabilities, provided future editions maintain comparable criteria or document any changes. Comparability from one year to the next cannot be assumed if the papers, materials, or rules change.
Information needed to interpret a comparison
These elements help distinguish a difference in capability from a difference in evaluation conditions.
| Element | What it clarifies |
|---|---|
| Papers and target results | What agents were asked to reproduce and how the cases were selected |
| Success criteria | What conditions a run had to meet to count as a reproduction |
| Budget and time limits | How many resources were allowed for each task and how the limits were applied |
| Agents, instructions, and environment | Which systems and conditions produced the reported rates |
| Logs, outputs, and evaluation | What evidence the evaluator saw and how the success criterion was applied |
A measure of capability with a defined scope
RECLAIM’s central contribution is to turn a broad task—reconstructing a research result—into a benchmark with goals and resources defined in advance. Its three tiers make visible how the work changes when code, data, and weights are available, when the model has to be trained, or when the method must be reimplemented. The published results show that, in this evaluation, reproduction became less frequent as fewer resources were available.
The strongest reading is also the narrowest: RECLAIM reports how four agents performed on 100 specific tasks under criteria and budgets set by the benchmark. Assessing the breadth of that evidence requires the full details of selection, evaluation, and execution. Assessing the science in the papers requires analyses beyond reproducing a single result.
Open questions
- The source information provided does not explain how the 100 papers were selected or identify the specific result set for each one.
- The GPU-hour amounts and time limits for individual tasks are not detailed.
- The abstract does not name the four agents or provide their individual results.
- The information provided does not confirm whether the environments, instructions, scripts, and complete evaluation rules are publicly available.
- The general success criterion is described as having been set in advance, but its operational rules for each paper are not included.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction