
One-sentence definition
Conjunto de tareas, condiciones y métricas usado para comparar un comportamiento concreto de uno o varios sistemas.
What is an AI benchmark?
An AI benchmark is a specified evaluation that combines a task or set of tasks, data, an execution protocol, and one or more metrics to describe how a system performs under those conditions. The term is also used for the evaluation package as a whole. In this article, benchmark refers to that evaluation design—not just the data or the number shown in a results table.
This practical definition matters because every score has a scope. For example, a score might indicate what proportion of a collection of problems a model solved using a particular execution method. It does not, without further evidence, show that the system is good at mathematics in general, answers every user correctly, or operates safely and reliably as part of a product.
Benchmarks are useful when you want to examine performance in a structured way, compare systems under shared conditions, or identify strengths and weaknesses. To understand what a result means, first establish which task it represents and which choices define the protocol. The word benchmark does not, by itself, guarantee that an evaluation is representative, fair, reproducible, or suitable for a particular decision.
How it works: tasks, data, protocol, and metrics
The task describes what the system is asked to do: answer a question, find information, modify code, or perform another bounded activity. The data are the specific cases used to test that task. A benchmark may also specify training or development examples, but these should be distinguished from cases held back for final evaluation. Using the held-out cases to tune the system can change what the test measures.
The protocol sets out how the evaluation is run. It may specify the input and response formats, available tools, instructions, time or compute limits, number of attempts, and scoring method. For systems that use tools, the environment and harness—the code that connects the model to the tasks, runs actions, and collects results—are part of the conditions. If these change, the result may change too, even if the model itself is unchanged.
A metric turns observed responses into a measure. Some metrics count correct answers; others assess qualities such as quality or safety. An aggregate result can summarize many cases in a single value, but an average can hide differences across task types and individual examples. Where available, check breakdowns, case counts, variability, and the rules used to handle ambiguous responses.
Finally, the model configuration identifies which system was evaluated and how it was run: its version, instructions, tools, and other relevant settings. A model label without these details is not enough to reconstruct the test. A clear methodological description helps readers interpret the result and judge whether another evaluation actually reproduces the same conditions.
A sequence for interpreting an evaluation
- 01Identify the task and the population of cases: what is being asked, and which situations are outside the test?
- 02Review the data, version, and any exclusion or selection rules.
- 03Check the protocol: configuration, tools, budget, and execution conditions.
- 04Read the metric and denominator: what counts as success, and how many cases is it calculated over?
- 05Before comparing results, check whether the conditions match and whether the cases resemble the intended use.
Three examples from different domains
These examples involve different tasks and different kinds of answers or evaluation criteria. They are not interchangeable, and none covers every capability of a system. The fact that an evaluation uses a well-known collection of problems or issues does not make its score a universal measure.
AIME 2024 belongs to the field of academic knowledge and mathematical problem-solving. The exam consists of mathematical problems, and its official solutions provide a reference for checking answers. If those problems are used to evaluate a system, the interpretation depends on which problems were included, the instructions given, whether tools were allowed, and how responses were judged. The result describes that protocol and that collection, not mathematical competence as a whole.
BrowseComp was created to evaluate web-browsing agents searching for information that is difficult to find. In an evaluation of this kind, the final answer is not the only consideration: browsing conditions, available tools, and how the information found is verified also matter. A good BrowseComp score does not prove that a system can find any web information correctly, regardless of source or update conditions.
SWE-bench focuses on real issues from software repositories: the system must propose changes that resolve problems described on GitHub. SWE-bench Verified is a selected and reviewed version of the benchmark. The project documentation describes a collection of 500 cases verified through human annotation and tests; it also notes that environment configuration, contamination, and the collection’s coverage limit the conclusions that can be drawn. A result therefore depends on the cases and harness used, and is not a guarantee that the system can maintain any codebase in production.
How to read a score—and when to compare scores
Before comparing two results, confirm that they refer to the same task, data version, split, metric, and scoring rules. Also check the model configuration, instructions, tools, budget, and environment. A higher number does not necessarily mean better performance if any of these elements changed. Even under similar conditions, differences may fall within execution variability or depend on a small number of cases.
The denominator helps put the evidence in perspective: a rate calculated over a few examples is not the same as one calculated over many. It is also useful to know which cases were excluded and how incomplete or ambiguous responses were handled. If a report provides only an aggregate score, look for results broken down by task or category before using it to choose a system.
Comparisons are more defensible when systems are run under a shared protocol and the details that could affect the result are documented. BetterBench examines issues such as purpose, scope, documentation, contamination, replicability, and comparability when assessing benchmarks. HELM, in turn, presents model evaluations through multiple scenarios and metrics and documents evaluation conditions. These approaches illustrate why a leaderboard without methodological context is incomplete evidence.
What to check before comparing two scores
| Element | Check | If it does not match |
|---|---|---|
| Task and data | Were the same cases and the same version evaluated? | The difference may reflect case selection or example difficulty. |
| Metric and aggregation | Is success counted in the same way and using the same denominator? | The figures may represent different things. |
| Model and execution | Do the version, instructions, tools, and budget match? | The difference cannot be attributed to the model alone. |
| Environment and judge | Do the execution environment and validation rules match? | Which responses are accepted, or whether a task can be completed, may change. |
Benchmark, dataset, metric, leaderboard, and custom evaluation
A dataset is a collection of data or examples. It can be part of a benchmark, but on its own it does not define the task, the full protocol, or how results are scored. The same collection can be used for different evaluations, so knowing the dataset’s name is not enough to know what was measured.
A metric is a rule for summarizing or judging results, such as an accuracy rate defined in a particular way. It is not the entire benchmark: different metrics applied to the same data can answer different questions. A leaderboard is a table or ranking system that presents participant results according to stated rules. It is a way of displaying scores, not proof that every entry was produced under comparable conditions.
A custom evaluation adapts cases and conditions to a specific need, such as the requests handled by an organization’s support assistant. It may resemble a benchmark, but it needs equally careful documentation: selection criteria, data, protocol, metrics, and limitations. An acceptance test, by contrast, checks agreed requirements for a product or system in a defined context. It may be part of an evaluation, but it is not automatically a general benchmark.
These distinctions help avoid shortcuts: a table with many scores is not necessarily a complete evaluation; a popular dataset is not necessarily representative of real use; and a familiar metric name does not, by itself, explain the success criterion.
Related terms, but not equivalents
| Term | What it refers to | What it does not establish on its own |
|---|---|---|
| Benchmark | An evaluation design: task, data, protocol, and scoring. | That it is representative, fair, or sufficient for a real-world decision. |
| Dataset | A collection of examples or data. | Which protocol or metric was used to evaluate it. |
| Metric | A rule for scoring or summarizing a result. | Which tasks were evaluated or whether the measure reflects intended use. |
| Leaderboard | An ordered presentation of results under certain rules. | That all results are comparable or have been independently reproduced. |
| Custom evaluation | A test designed for a particular use case or population. | That its results generalize beyond that context. |
Limitations: contamination, overfitting, and external validity
Contamination occurs when information from evaluation cases—or very similar answers—appears in the data used to train or tune a system. In that situation, a correct response might reflect prior familiarity with the material rather than only the capability the evaluation was intended to measure. Contamination is not always easy to detect: training data may not be public, and exact-text matching will not reveal every kind of overlap.
Benchmark overfitting can occur when developers repeatedly optimize for strong results on a known test. That adaptation may improve the score without improving performance on new tasks to the same extent. Holding back evaluation sets, documenting changes, and supplementing the test with different cases can help reduce this risk, but cannot eliminate it completely.
External validity asks how far a result can inform us about situations beyond the benchmark: other users, domains, languages, tools, software versions, or production conditions. A controlled collection makes comparison easier, but may omit important real-world requirements such as latency, cost, privacy, extended interaction, recovery from errors, or human oversight.
Execution can also vary. A system may produce different responses across attempts; environments may fail; and, when automated or human judges are involved, their rules and disagreements can affect the result. It is useful to publish the evaluation method, instructions, and environment, as well as examples of inputs and outputs where possible. A single rounded figure can conceal uncertainty, uneven results across subgroups, or serious failures on specific cases.
Aggregate scores simplify interpretation, but they compress choices: which tasks count, how much each one weighs, and how errors are treated. A high average can coexist with weak performance in a relevant category. For practical decisions, examine the distribution of results and the failures with the greatest potential impact, not just the overall rank.
When a benchmark is useful—and when it is not
A benchmark can help compare systems under stated conditions, track changes between versions, identify weak areas, and establish a shared reference for a bounded task. It can also help frame follow-up questions: which kinds of cases account for the errors, what resources were needed, and whether an improvement appears across more than one evaluation.
A benchmark is not enough if it is the sole argument for deploying a system, predicting its quality in an unrepresented context, or claiming that a capability is general. Those decisions call for complementary evaluations: tests using relevant data and users, failure analysis, safety and privacy checks, and operational measurements such as cost or latency, depending on the purpose.
Before adopting a benchmark, define the decision it is meant to inform. If the question is specific—for example, whether an assistant correctly classifies incidents for a service—a well-designed and documented custom evaluation may be more relevant than a public ranking. It can be used alongside established benchmarks, but should not be confused with them.
Practical criteria for using a benchmark
- 01State the decision question before choosing a test.
- 02Check that the tasks and cases resemble the use you care about.
- 03Read the protocol, metric, version, configuration, and number of examples.
- 04Compare results only when the conditions are sufficiently equivalent.
- 05Review errors and breakdowns; do not rely on a single aggregate score.
- 06Supplement the benchmark with a use-case evaluation and state the uncertainties.
Related concepts
To explore the topic further, see the glossary entries for benchmark, evaluation, reproducibility, and data contamination, along with the main glossary. These concepts help clarify what was measured, whether other people can repeat the test, and whether a result may be inflated by prior exposure to the cases.
The practical takeaway is simple: ask what task was performed, using which data, under what protocol, and according to what criterion. If those pieces are unclear, the score is not a sufficient basis for comparing systems or anticipating their performance in a different setting.