The question: does the agent understand what caused a change?
A research agent can run code, choose parameters, and report a score. But those actions alone do not show whether it understands how changing a component affects the outcome. WhatWorkedBench proposes measuring that ability through predictions about experimental changes: given a workflow and its options, can the agent anticipate the results it would obtain under different configurations?
The work presents itself as a benchmark for experimental understanding. Rather than evaluating only whether an agent completes a task or reaches a high score, it asks the agent to produce predictions for the possible configurations of a workflow. The central measure is therefore the quality of its predictions about the effects of changes, compared against reference results.
The distinction matters. A high score on one configuration does not necessarily show that an agent can explain which component produced it, or what would happen if that component changed. Conversely, an agent might approximate effects without identifying the optimal configuration under a limited budget of new measurements. These are related, but not identical, capabilities.
How the evaluation works
According to the preprint abstract, agents inspect code, choose which measurements to make within a budget, and submit a response surface: a table predicting the score for every configuration of the components. The task therefore does not end with running a single experiment. The agent must extrapolate from the observations available and represent how results would vary.
To build the references, the authors exhaustively run the configurations on CPU. They then calculate the effect of changing each component while holding the others fixed. The abstract also says the analysis accounts for combinations of changes across components. This reference makes it possible to compare predictions with observed results across the set of configurations.
This procedure provides a quantitative basis for the benchmark, but it depends on the exhaustive runs and the defined configurations adequately representing each task. The comparison speaks to those specific workflows and options; it does not remove the uncertainty involved in transferring conclusions to other code, objectives, or experimental contexts.
The steps in an evaluation
- 01The agent inspects the workflow code and its configurable components.
- 02It selects new measurements within the assigned experimentation budget.
- 03It predicts scores for possible configurations in a response surface.
- 04The predictions are compared with reference effects calculated from exhaustive runs.
Scale and composition reported in the preprint
The preprint abstract reports 36 tasks, built from 30 data sources and distributed across eight workflow types. It also describes 1,248 configuration records. For the main evaluation, it reports 4,206 numerical-control records spanning all eight families and 108 agent episodes across the original six families.
These figures describe different units: tasks, data sources, workflow types, configurations, control records, and episodes are not interchangeable quantities. In particular, the episode count should not automatically be read as the number of agents, nor should control records be treated as independent experiments run by agents. The abstract does not provide, in the material available here, a complete breakdown of each figure by task.
The project's public repository is presented as a complementary source for examining tasks, the evaluator, numerical controls, recorded episodes, guides, and tests. This makes it easier to inspect the structure of the work and look for materials that could support reproduction. However, the existence of a repository alone does not establish that every figure can be reproduced without additional dependencies, data, and execution conditions.
What each figure in the abstract represents
| Item | Reported count | Cautious interpretation |
|---|---|---|
| Tasks | 36 | Evaluation cases; not the same as 36 workflow types. |
| Data sources | 30 | Sources associated with the task set. |
| Workflow types | 8 | Families of experimental procedures. |
| Configuration records | 1,248 | Recorded configurations, not agent episodes. |
| Numerical controls | 4,206 | Control records reported across all eight families. |
| Agent episodes | 108 | Episodes from the original six families, according to the abstract. |
Reported results and how to interpret them
The abstract reports that, with eight new measurements, a method called pair-effect ridge selected an optimal configuration on 15 of 22 sources. It also kept every effect error within 10% of the score range on three sources. These are results for specific subsets and conditions; they do not mean that the method found the optimum on every source or achieved that error bound in general.
The work also compares predictions fitted with a Gaussian process to observations collected by agents. In the original Flash cohort, effect recovery rises from 0.632 to 0.698; in an additional cohort, it rises from 0.621 to 0.720. The abstract also presents an analysis of six completed submissions on beat-detection and graph tasks, where family-macro recovery increases from 0.303 to 0.455 when the same type of model is fitted to the agent observations.
In another analysis, covering six workflows with six binary options and a budget of 20 new measurements, encoding code equivalences—configurations with identical behavior—raises Gaussian-process recovery from 0.248 to 0.462. A reasonable interpretation is that using known program structure can improve predictions in that scenario. These figures do not establish that the same increase will occur in other workflows or under other budgets.
The comparisons described depend on the metrics, cohorts, and episodes specified by the study. The figures are evidence about that experimental design, not a universal ranking of agents. Assessing differences between systems would also require details about which models were evaluated, how they were selected, and how tasks were allocated.
What remains to assess about models and reproducibility
The available abstract gives aggregate figures and names some methods, but it does not identify here every agent and model evaluated or provide their full comparative results. Nor is it enough to reconstruct the parameters, partitions, execution conditions, or exact steps of every analysis. The full preprint is identified as the source for the protocol and models, while the repository offers implementation and reproduction materials according to the description provided.
The statement that references come from exhaustive CPU runs describes how the reference effects were obtained. Reproducing them would require checking the code and records to see which configurations were run, how the metrics were computed, and what data or dependencies each task requires. An exhaustive run over a defined configuration space should not be assumed to cover every possible intervention in a real scientific problem.
The work is identified as an arXiv v1 preprint. The information provided does not document peer review. It is therefore more precise to treat it as preliminary research made publicly available, rather than a result already validated through peer-reviewed publication. This status does not invalidate the benchmark, but it matters when calibrating confidence and looking for independent confirmation.
Conclusion: a bounded tool for studying experimental agents
WhatWorkedBench addresses a specific and useful question: can an agent anticipate the effects of changes to components in experimental workflows after making a limited number of measurements? Its protocol turns that question into predictions that can be compared with references obtained through exhaustive runs, and the work reports improvements associated with fitting methods and the use of code equivalences in particular scenarios.
The conclusion should remain tied to the benchmark's scope. Predicting a response surface for defined tasks is not the same as formulating scientific hypotheses, choosing relevant problems, recognizing spurious results, or conducting research autonomously. Nor does it predict performance on workflows not represented by the evaluated tasks.
For readers following AI system evaluations, the main contribution is a more specific way to ask what an agent can do: not just whether it obtains a result, but whether it can predict how that result will vary when parts of the procedure change. The reported figures are promising in some analyses, but full model comparisons, independent reproduction, and review of the work remain important for determining how far the conclusions extend.
Open questions
- The available abstract does not identify every evaluated model and agent or present the full comparison between them.
- The information available here is not detailed enough to reconstruct the exact parameters, partitions, dependencies, and execution conditions for every result.
- The public availability of repository materials does not, by itself, confirm independent reproduction of all figures.
- Exhaustive runs establish references for the defined configurations, but do not necessarily cover every possible intervention in real research.
- The information provided does not indicate that the preprint has been peer-reviewed.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction