The historical question: What does it mean to evaluate an AI capability?
Evaluating an artificial intelligence capability means turning a broad question—for example, whether a system understands instructions or can help solve problems—into a test with tasks, conditions, and a rule for judging the results. The score produced does not measure capability in the abstract. It summarizes performance under those specific conditions.
A correct answer to a question, a solution that passes software tests, and an operation that leaves an application in the requested state are different kinds of evidence. They are not interchangeable units. Each makes certain successes and failures visible while leaving others out. The history of evaluation can therefore be read, in part, as a change in what counts as the unit being evaluated: first, answers to bounded examples; then, broader collections of tasks; and, in some recent benchmarks, sequences of actions executed in environments.
This sequence is an interpretive framework, not an exhaustive chronology or an inevitable ladder toward a better measure. The selected studies represent different approaches. An answer-based benchmark may be appropriate for a narrowly defined question; an interactive one may offer more direct evidence about execution, but it also adds dependencies on the environment and protocol.
The first stage: Bounded tasks and scorable answers
In a dataset-based evaluation, each example provides an input and specifies which output counts as correct, or which criterion should be applied. In language understanding, the input might be a sentence or a pair of sentences; the output might be a label, a judgment, or an answer. The aggregate result makes it possible to compare systems that took part in the same task under a common protocol.
This design has practical advantages. Examples can be repeated, results can be calculated consistently, and researchers can compare methods without asking them to operate a complete application. If the question is whether a model can distinguish a particular semantic relationship, a well-defined classification task can provide useful evidence.
But the unit being evaluated is narrow. Producing the correct label does not demonstrate that a system can plan a series of steps, use tools, react to an unexpected change, or complete a task in an interface. Nor does an aggregate score, by itself, explain where errors are concentrated. Data coverage, the way questions are framed, and the chosen metric define the limits of what can be inferred.
The cautious inference is conditional: the system achieved a certain result on those tasks, with that data and that scoring rule. Supporting broader claims about general competence, robustness, or usefulness in a real-world situation requires additional tests.
Broadening the test field: GLUE and BIG-bench
Introduced in 2018, GLUE brings together several natural language understanding tasks in a multitask benchmark and analysis platform. Its design shifts the focus from a single test to a set of related problems: it makes it possible to observe whether a system performs across different tasks and provides a way to summarize some of that performance. The range of tasks makes it harder to reduce evaluation to a single skill, although it does not remove the limitations of individual tasks or turn the aggregate result into a universal measure.
BIG-bench broadened the range of tests assembled even further. Rather than focusing on a relatively bounded family of language problems, it proposed a large collection of tasks contributed by different collaborators, with varied goals and evaluation methods. The paper examines how language models perform across that collection and how results change as model scale varies.
The important change is not simply an increase in the number of examples. A heterogeneous collection can test different capabilities and behaviors, and reveal that a system that is strong on one task is not necessarily strong on another. At the same time, that variety makes comparison more complicated: tasks may have different formats, metrics, and difficulty levels. An aggregate figure makes it easier to get an overview, but it can conceal important differences between components.
In both cases, evaluation remains primarily a test of answers to tasks defined in advance. Having many tasks is not the same as observing an agent working in an open environment. Breadth improves coverage within the collection; it does not guarantee that the collection represents every possible use or measures extended execution.
What each approach makes observable
| Approach | Unit evaluated | Question it helps answer | Main limitation |
|---|---|---|---|
| Bounded dataset | Answer to an example | Did it get this task right under this metric? | Does not, by itself, test execution across a sequence of actions. |
| Multitask benchmark such as GLUE | Results across several tasks in one family | How is performance distributed across related problems? | The aggregate can hide differences between tasks. |
| Diverse collection such as BIG-bench | Answers across a broad range of tasks | What patterns emerge as tasks and models vary? | Metrics and conditions may differ from task to task. |
| Executable or interactive task | Actions and final state in an environment | Did the system produce the expected operational result? | The result also depends on the environment and verifier. |
Changing the unit: Resolving an issue in a repository
SWE-bench shifts the unit of evaluation toward a software engineering task: resolving issues from real GitHub repositories. Rather than judging only a text response to a question, the system has to work with the context of a project and produce code changes that address the stated problem.
The evaluation can check the patch using the project’s tests, including tests related to the reported bug and tests that should continue to pass. This provides evidence about something more concrete than a convincing explanation: whether the proposed change works according to the checks available in the benchmark environment. The evaluator can determine whether the result passes those tests; that is not the same as certifying that the patch is the only correct solution, that it is well designed in every respect, or that it is safe for any deployment.
The evaluated system is not necessarily just the model in isolation. Depending on the configuration, the result may depend on how the repository is presented, which tools are available for inspecting or editing files, how many attempts are allowed, and how the tests are run. To compare scores, it is important to know which components are included and whether the protocol was held constant.
The evaluation still consists of a defined collection of issues and execution conditions. Therefore, resolving an issue demonstrates success on that case under its checks, not the ability to solve every software problem. A repository-based test brings the task closer to a specific professional activity, but it does not automatically cover product requirements, collaboration, long-term maintenance, or the consequences of a change in production.
Adding interaction and state: Tasks in computer environments
OSWorld evaluates multimodal agents on open-ended tasks in real computer environments that are simulated. Instead of merely producing an answer about an application, an agent may have to observe the interface and act through controls such as a keyboard or mouse to reach a requested state. The evaluation focuses on tasks performed in the environment, not just on the linguistic quality of a description.
This shift makes it possible to observe dimensions that an answer-based test does not directly record: whether actions are executed, whether a sequence reaches the intended result, and whether the final state meets an evaluation condition. It also makes the relationship between perception, decision, and action more visible. A system may correctly describe what it should do and still fail to do it; in an interactive environment, that difference can be part of the result.
Execution provides a signal closer to the operational task, but it does not remove the need to decide what counts as success. The environment, applications, initial state, instructions, tools, and verifier all form part of the test conditions. If the evaluation compares states using automated rules, those rules can check specific aspects of the result; they do not necessarily judge every detail of the quality, safety, or appropriateness of the path taken.
The agent also needs to be distinguished from the environment. A score obtained with particular tools and controls does not automatically describe what the model would do without them, with a different interface, or under conditions not covered by the test. The result belongs to the system and protocol that were evaluated, not to an isolated capability of all those components.
How to read an executable evaluation
- 01Identify the task and initial state: what must be achieved, and from what starting situation?
- 02Record which system is being evaluated: the model, tools, control interface, and execution limits.
- 03Check how progress is observed: through action logs, application state, or both.
- 04Read the success rule: what condition does the evaluator verify, and which aspects does it not inspect?
- 05Limit the conclusion to the observed result and the conditions described.
What changes with an executable environment—and what remains
The move from scored answers to executable tasks broadens the observable evidence. It can show whether the system produces a change in a repository or an application, rather than merely writing a plausible solution. That is a meaningful difference for claims about the ability to act. However, it does not automatically turn a benchmark into a complete measure of usefulness, autonomy, or reliability.
It is useful to separate four elements. The task defines what is requested; the environment determines where and under what conditions it happens; the verifier establishes which results it considers satisfactory; and the evaluated system includes the components that receive the task and produce actions. A change in any one of these elements can change the difficulty or meaning of the score. If results from different protocols are compared without acknowledging those changes, some of what is attributed to the model may actually be due to other conditions.
The validity of the verifier deserves particular attention. An automated test can be reproducible and useful, but it evaluates only what it checks. An incomplete test suite might fail to detect an uncovered defect; a check of the final state might not penalize an inefficient path if efficiency is not part of the criterion. These are general possibilities to investigate for each benchmark, not flaws that should be attributed to a particular evaluation without evidence.
Reproducibility also depends on operational details: software versions, initial state, access to tools, time budget, or number of attempts. A description of these limits makes the result easier to interpret. When details are missing, comparison may be less informative; in that case, it is better to state the uncertainty than to fill the gaps with assumptions.
Decision guide: What evidence supports the claim?
| Claim to be supported | Relevant evidence | Caution |
|---|---|---|
| The system solves this type of question | Results on comparable answer-based tasks under a described protocol | Do not automatically extrapolate to interaction or execution. |
| Performance covers several related tasks | Disaggregated and aggregate results from a multitask benchmark | Check which tasks make up the aggregate and how they are weighted. |
| The system changes code to address issues | Execution of code changes and tests associated with the issues | Passing tests does not prove that every possible requirement is met. |
| The system completes actions in applications | Executed tasks, resulting state, and evaluation rules | The conclusion depends on the environment, controls, and verifier. |
| The system is reliable or autonomous in everyday use | Additional evidence under varied conditions relevant to that use | A single benchmark score is not enough to support this conclusion. |
How to read a historical score
A historical score needs context. The first step is to identify which version of the benchmark and which protocol were used. The same label may refer to different task collections, metrics, or configurations. Before comparing two figures, confirm that they describe tests that are sufficiently comparable.
The second step is to identify the unit of evaluation. Was the system scored on one answer, a collection of answers, a patch tested against software checks, or a task in an interface? That difference determines what kind of evidence the result offers. A percentage of correct answers and a task-completion rate are not equivalent scales, even if both are expressed as percentages.
Next, separate the aggregate result from its breakdown. In a multitask collection, looking at performance by task can reveal strengths and weaknesses that disappear in the average. In interactive evaluations, it matters which execution conditions were held constant and which states the verifier considered satisfactory. When reports make disaggregated results available, they help prevent a summary figure from becoming a broader claim than the evidence allows.
Finally, an improvement across evaluations does not, by itself, show how much a general capability has advanced. If the model, task collection, tools, or scoring rules all change at once, the entire change cannot be attributed to a single cause without further analysis. A comparison is stronger when conditions are held constant or when differences are described precisely.
Conclusion: Choose evidence to match the claim
GLUE and BIG-bench show two ways to broaden answer-centered evaluations: assemble related tasks or cover a more diverse collection. SWE-bench and OSWorld illustrate approaches that shift testing toward executed results in repositories and computer environments. They are not interchangeable rungs in a ranking; they answer different questions and leave different kinds of evidence.
When a claim concerns answers to language tasks, a relevant dataset can be a useful test. If the aim is to understand how performance is distributed across several tasks, it matters to examine a multitask evaluation and its disaggregated results. If the claim concerns changing a project or performing actions in an application, an executable evaluation can provide more direct evidence about those operational results.
The final rule is simple: interpret the score at the level of the test that produced it. An evaluation closer to real use may reveal more about actions and states, but it does not, by itself, measure every aspect of a useful, safe, or reliable system. Supporting those conclusions requires transparent protocols and additional evidence matched to each claim.
Open questions
- The task coverage of any benchmark does not, by itself, allow conclusions about performance in every real-world use.
- Automated tests can verify specific conditions without proving that a solution is complete, optimal, or safe in every respect.
- Historical comparisons can be ambiguous when versions, tools, budgets, or evaluation rules change.
- The cited benchmarks are representative examples of different approaches, not an exhaustive chronology of AI evaluation.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction