Ilustración editorial para RECLAIM mide si los agentes de IA pueden reproducir resultados de investigación
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A reproduction test, not an isolated programming task

RECLAIM is a benchmark designed to measure whether an artificial intelligence agent can reproduce a specific result described in a machine learning paper. The preprint presents a set of 100 NeurIPS 2025 papers and defines one task for each: using the paper and any materials released by its authors, the agent must work within a GPU-hour budget set in advance.

The scope of the work distinguishes this from a conventional programming test. To reach a result, an agent may need to install software, resolve errors, understand the method, run experiments, and check its results. The evaluation aims to cover that chain of tasks, rather than simply determine whether the system produces code that looks correct.

The benchmark also specifies in advance which result must be reproduced and what conditions count as success. That is an important choice: without a predefined goal, comparisons between agents could depend on differing judgments about what constitutes a satisfactory reproduction. The preprint’s abstract does not detail the exact rules used for each paper, so it is not enough to reconstruct how the criterion was applied in every case.

02

Three tiers based on what the authors released

RECLAIM groups tasks according to the availability of resources. In the Run tier, the code, data, and model weights are available. In Retrain, the weights are missing, so the agent must train the model. In Reimplement, the code is unavailable and the agent must write an implementation. According to the preprint, the materials released by the authors determine the difficulty tier.

This classification helps put the results in context: tasks do not all start from the same point. Running an already prepared system, rebuilding its weights through training, and implementing a method from scratch require different kinds of work. A single overall success rate, without a breakdown by tier, would therefore hide important differences.

The classification does not mean that every task within a tier has identical requirements. Papers may describe different methods and experiments. The available abstract does not list the resources for each paper or explain how much their requirements vary. The tiers therefore describe the availability of certain materials, not a perfect equivalence in difficulty.

How to read the RECLAIM tiers

The table summarizes the three tier definitions given in the preprint. It is not an independent ranking of the difficulty of individual papers.

TierResources describedWork the agent must take on
RunCode, data, and weightsRun the released materials and obtain the specified result
RetrainWeights are missingTrain the model, in addition to completing the rest of the task
ReimplementCode is missingWrite an implementation of the method in an attempt to reproduce the result
03

The results show a gap between tiers

The preprint reports that four agents were tested once on each paper. The best agent in each tier reproduced 41% of Run papers, 27% of Retrain papers, and 15% of Reimplement papers. In this set of tasks, and under the described procedure, success declined as resources needed to rebuild the system were removed.

These figures are benchmark results, not a universal measure of the capabilities of any agent. Nor should they be read as a direct comparison between agents without further information: the abstract reports the best result in each tier, but does not give the four systems’ names, their individual scores, or enough detail to reconstruct the variability across attempts.

The paper also reports that failed attempts used, on average, 29% of the assigned budget. The preprint’s interpretation is that many agents stopped with budget still available. This suggests that the compute limit alone does not explain every failure; by itself, however, the figure does not establish why each attempt stopped or what change would have been enough to achieve a reproduction.

Another frequent error was implementing the method without checking any part of it against the values reported in the paper. The abstract records this error in 63 of 400 runs. This observation points to a distinction between producing a plausible implementation and checking it against evidence: writing out the method does not guarantee that the agent reproduced the relevant conditions.

A cautious way to read the success rates

  1. 01Identify the resource tier: Run, Retrain, or Reimplement.
  2. 02Read the rate as the best agent’s result in that tier, not as the average across all agents.
  3. 03Keep in mind that the report describes one run per agent and paper.
  4. 04Do not turn the percentage of reproduced results into a claim about the overall validity of the papers.
04

What it evaluates—and what it leaves out

RECLAIM evaluates whether an agent can reach a previously selected result using the available materials and within a computational budget. According to the abstract, a separate instance of a language model grades the runs based on logs and outputs, rather than relying on the agents’ written reports. The goal is to assess what happened during execution, not merely what the system says it did.

That also defines the limits of the conclusion. Reproducing a specific result does not automatically verify every methodological choice in a paper, the quality of its data, the robustness of its analyses, or the validity of its scientific conclusions. Conversely, failure to complete a benchmark task does not by itself show that the original result is wrong: there may be technical, implementation, or resource-related causes that the abstract does not break down.

This distinction matters for readers and teams that want to use the benchmark as a signal of progress. A low success rate may reveal that agents struggle to reconstruct experiments from incomplete information; without additional evidence, it should not be treated as a verdict on the paper being evaluated.

05

Reproducing the evaluation also requires details

To compare agents independently, knowing the benchmark’s size and success rates is not enough. Researchers would need access to the paper selection, the target result for each task, the operational success criterion, the specific GPU budgets, and the time limits. The agents and their configurations, the instructions they received, the environments, the available materials, and the logs or outputs used to grade each run also matter.

The preprint abstract confirms that RECLAIM sets the target result, success criterion, and GPU-hour budget for each paper, and that evaluation is based on logs and outputs. However, the information provided here does not specify the budget values, how the 100 papers were selected, the names of the four agents, or whether the environments and scripts needed to repeat the evaluation have been published. Those points should be checked in the paper and its accompanying materials before attempting a more detailed comparison.

The presentation describes RECLAIM as a benchmark that can be rebuilt each year from new conferences. This offers a possible way to track changes in agent capabilities, provided future editions maintain comparable criteria or document any changes. Comparability from one year to the next cannot be assumed if the papers, materials, or rules change.

Information needed to interpret a comparison

These elements help distinguish a difference in capability from a difference in evaluation conditions.

ElementWhat it clarifies
Papers and target resultsWhat agents were asked to reproduce and how the cases were selected
Success criteriaWhat conditions a run had to meet to count as a reproduction
Budget and time limitsHow many resources were allowed for each task and how the limits were applied
Agents, instructions, and environmentWhich systems and conditions produced the reported rates
Logs, outputs, and evaluationWhat evidence the evaluator saw and how the success criterion was applied
06

A measure of capability with a defined scope

RECLAIM’s central contribution is to turn a broad task—reconstructing a research result—into a benchmark with goals and resources defined in advance. Its three tiers make visible how the work changes when code, data, and weights are available, when the model has to be trained, or when the method must be reimplemented. The published results show that, in this evaluation, reproduction became less frequent as fewer resources were available.

The strongest reading is also the narrowest: RECLAIM reports how four agents performed on 100 specific tasks under criteria and budgets set by the benchmark. Assessing the breadth of that evidence requires the full details of selection, evaluation, and execution. Assessing the science in the papers requires analyses beyond reproducing a single result.

Open questions

  • The source information provided does not explain how the 100 papers were selected or identify the specific result set for each one.
  • The GPU-hour amounts and time limits for individual tasks are not detailed.
  • The abstract does not name the four agents or provide their individual results.
  • The information provided does not confirm whether the environments, instructions, scripts, and complete evaluation rules are publicly available.
  • The general success criterion is described as having been set in advance, but its operational rules for each paper are not included.
07

Keep exploring

07

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction