The key question: what does it mean for a benchmark to be contaminated?
Saying that an AI benchmark is contaminated can describe several different situations. An evaluation question might appear verbatim in data used to train a model; it might recur with formatting changes; or the model might have learned patterns during training that are very similar to those required by the test. These possibilities do not carry the same evidentiary weight, nor do they automatically imply the same effect on a score.
It helps to distinguish three claims. The first is that there is overlap between evaluation material and some relevant dataset. The second is that the model was exposed to that material before it was evaluated. The third, and more demanding, is that the exposure caused an improvement in the result. A detector may provide evidence relevant to the first questions; demonstrating the third usually requires comparing results under conditions that allow the effect to be isolated.
This distinction matters when interpreting any warning. A textual match can point to potential exposure, but it does not reveal when the text entered the data, whether the model learned it, or how much it influenced the response. Conversely, finding no exact match does not rule out paraphrases, derived data, or familiarity with the type of task.
This article is limited to the studies verified and listed at the end. They help explain several methods and their assumptions, but they are not enough to verify the aggregate figures, study count, or coverage date attributed to a systematic review mentioned in the editorial proposal. Accordingly, neither a corpus of 55 studies nor a general inflation range of 6% to 40% is presented here as a verified fact.
An operational distinction between four forms of contamination
When examining a specific allegation, it is useful to classify what overlaps. This classification is a practical analytical framework, not a claim that all the studies reviewed here share a single taxonomy. The categories can also occur together: a task may have circulated verbatim while also representing a problem pattern that is widespread in the training data.
Exact matching is the most visible case: a question, reference answer, or solution appears without substantial changes in training material. A text search can find matches, but their interpretation depends on which corpus could be inspected, how complete it is, and what transformations were applied to the data. The absence of a match in an incomplete corpus does not prove the absence of exposure.
Surface-level transformations include changes in formatting, punctuation, order, or wording that preserve much of the content. A literal comparison may miss them. Searching for fragments, normalizing text, or using other retrieval procedures may uncover candidates, although these approaches can also produce false positives: common expressions or task templates may appear in many documents without any copy of the evaluation item being present.
Semantic equivalence is harder to establish. Two prompts may call for the same reasoning in different words, without sharing long sequences of text. Semantic similarity can suggest possible pairs, but it does not prove that the model saw either one or that the exposure explains its response. Determining whether two examples are equivalent may also depend on human annotations and criteria.
Exposure to a task type is a different claim from copying examples. A model may have received many problems, questions, or code snippets in the same general format without any specific benchmark instance being identified in its data. That familiarity might make it easier to answer, but measuring it requires defining what counts as exposure and separating its effect from general capabilities acquired during training.
Five method families and what they require
No single detector can answer every question. Methods observe different signals and depend on different levels of access. The table below organizes five approaches that are useful when reading contamination studies; it is a comparative guide, not a universal reliability ranking. In particular, a more direct signal of a match does not necessarily provide a better estimate of the causal effect on a score.
What each approach observes—and what it leaves unresolved
| Family | Typical access or input | What it can indicate | Main limitation |
|---|---|---|---|
| Direct text matching | Benchmark items and accessible training corpus, or a candidate collection | The presence of identical or very similar sequences in the material inspected | Corpus coverage conditions the result; a match does not prove the model used the material or that it affected the score |
| Document retrieval | Test items and access to a corpus to search; may rely on text retrieval | Candidate documents containing content related to the items | Results depend on the collection, query, and thresholds; retrieving a document does not by itself establish that it was part of training |
| Testset Slot Guessing | Access to a model that allows probes on parts of examples, according to the protocol | Signals that the model can complete or recognize elements of a test | A correct answer may reflect general capability or knowledge of the format, not necessarily exposure to the specific instance |
| Membership inference | Access to model responses or signals and to examples for comparison, depending on the method | A statistical difference consistent with an example having been part of the training data | Validity depends on assumptions about training and the data distribution; available evaluations show limits to reliability |
| Performance comparison | Model results on a primary benchmark and reference tests, as well as explicit assumptions | A performance pattern consistent with contamination and a possible performance difference | The comparison depends on the choice of references and the assumptions; it does not necessarily identify the cause of the performance |
What the available studies allow us to claim
A study published at NAACL in 2024 presents corpus-retrieval methods and a protocol called Testset Slot Guessing for investigating benchmark contamination, including settings involving open and proprietary models. These methods extend the options beyond searching for literal copies, but the evidence they produce still depends on how queries are defined, which corpora are available, and what access the model allows.
A survey and evaluation published in Findings of NAACL in 2025 examines detector assumptions and evaluates membership-inference methods. Its most relevant finding for a critical reading is that some tests performed close to chance on corpora used for pretraining. This does not mean that all membership inference is useless; it does warn that calling a method a “detector” is not enough to make its classification conclusive.
ConStat proposes performance-based detection by comparing results on primary and reference benchmarks under defined assumptions. The comparison may help estimate whether observed behavior is consistent with contamination, but it does not automatically turn a difference into causal proof. To assess it, one must examine which references were chosen, why they are comparable, and what alternative explanations could produce the same difference.
A controlled study on machine translation deliberately changes exposure to evaluation data to measure changes in model performance. The paper reports that effects vary depending on whether the input, the output, or both are contaminated. This design is informative because it manipulates exposure rather than inferring it from observed results alone. Even so, its conclusions apply to the experimental conditions studied: they cannot simply be transferred to other benchmarks, tasks, or training processes.
Taken together, these studies show why it is helpful to ask narrowly framed questions. Was a string found in a particular corpus? Was a statistical signal consistent with membership detected? Were correct responses to probes observed? Or was a change in performance measured after introducing controlled contamination? These are different results. Describing them with the same label can conceal important differences in access, assumptions, and inferential strength.
Why inflation estimates are not interchangeable
An estimate of how much contamination raises a score can only be interpreted alongside the procedure that produced it. The benchmark, the unit considered contaminated, the proportion of affected examples, the task, the model, and the training stage all matter. So do the reference data and the comparison design. Two percentages can describe different quantities even if both are called “inflation.”
For example, one estimate might compare performance with and without exposure introduced under controlled conditions. Another might infer possible overlap in corpora and project its effect using assumptions. A third might compare scores from different models or sets of items. Without sufficient information about these designs, it is not rigorous to rank their figures as if they were repeated measurements of the same quantity.
The editorial proposal mentions an approximate range of 6% to 40% associated with published estimates for MMLU, GSM8K, HumanEval, and HellaSwag. The verified sources available for this article do not establish which studies support each endpoint, what definition of inflation was used, or whether the figures are comparable. For that reason, the range is not presented here as a confirmed result or as a common measure of contamination across benchmarks.
The same caution applies to benchmark names and aggregate figures from reviews. Before repeating an estimate, it is worth returning to the primary study and recording what was measured: the percentage of items with matches, the number of examples retrieved, the score difference, the proportion of correct answers attributable to an experimental condition, or some other variable. If the unit is unclear, the figure can give an impression of precision that the design does not support.
Presence, exposure, and causality: an evidence ladder
A sound analysis avoids jumping from a signal to a causal conclusion. First, describe the finding precisely: for example, a search identified similar documents in a particular collection. Next, assess whether that collection reliably represents the relevant training data and whether it is possible to establish that the model had access to the documents before evaluation. Finally, attributing a score change to that exposure requires a comparison that convincingly rules out alternative explanations.
Observable model memory is not identical to its training history either. A probe may prompt the model to complete part of an example, but that success could depend on general knowledge, task regularities, or clues included in the prompt. Conversely, the fact that a model does not reproduce an item does not prove that it never saw it: retrieval can fail, especially if the query changes or the model does not produce a literal reproduction.
Causal inferences are clearer in a controlled experiment when exposure is manipulated and performance is measured under comparable conditions. Even then, the conclusion is local to the protocol. It is not enough for an experiment to show that an effect is possible in order to attribute a particular model’s score to contamination; the experiment would need to be connected to the actual exposure, model, data, and benchmark under investigation.
It is therefore useful to describe evidentiary strength in graduated terms: a candidate overlap, evidence of probable exposure, or an experimental demonstration of an effect under specified conditions. These are not universal official labels, but a way to make clear what has been observed and what inference remains open.
Practical questions for reviewing a contamination claim
When evaluating a claim, the first step is to locate the original study and distinguish its findings from third-party interpretations. A press release, comparison table, or synthesis can be useful as a guide, but the strength of the claim depends on the primary work’s data, protocol, and limitations.
It is also worth checking what access the team had. Did it inspect the training data, or only an approximate corpus? Did it have access to probabilities, weights, or only model responses? Is the method reproducible for a closed model? Lack of direct access does not automatically invalidate an analysis, but it changes what can be concluded and how much the conclusion depends on assumptions.
The following checklist can help both those who publish evaluations and those who interpret them. If several questions remain unanswered, the appropriate response is to describe the conclusion as provisional, specify the kind of signal observed, and avoid attributing any score difference to contamination.
Checklist for interpreting results
- 01Define the object of the claim: is it about textual overlap, probable exposure, observable memory, or an effect on the score?
- 02Identify the data compared and the coverage of the corpora; state what material could not be examined.
- 03Describe the access required: corpus, model responses, probabilities, or parameters, as applicable.
- 04Review possible false positives and false negatives, as well as the thresholds and criteria used to consider two examples similar.
- 05Look for an appropriate comparison to estimate the effect: what alternative condition indicates what would have happened without exposure?
- 06Check whether training stages and types of exposure are distinguished; do not assume they are equivalent without evidence.
- 07Verify that the reported figure has an explicit definition, denominator, and scope.
- 08Match the conclusion to the design: state what was observed, what is inferred, and what remains uncertain.
Conclusion: treat contamination as a risk that requires graduated evidence
Benchmark contamination is an important issue when interpreting evaluations, but it is not an automatic explanation for a high score or a property that any single detector can resolve definitively. A textual match, a membership signal, a correct response to a probe, and an effect measured in an experiment are different kinds of evidence.
The studies available here offer complementary methods and also document limitations: retrieval depends on the data inspected, membership inference can be unreliable under certain conditions, and performance-based estimates require assumptions. Controlled experiments help investigate causality, but their results should be limited to the conditions they actually tested.
The most transparent practice is to report the method, access, comparison decisions, possible errors, and uncertainty, and to distinguish exposure explicitly from effect. A transparency card may be useful as a documentation proposal, but the verified sources used here do not establish the content of a specific card or justify presenting it as an adopted standard. For readers and researchers, the central principle remains simple: require each conclusion to match the evidence supporting it, without turning suspicion into a demonstrated cause.
Open questions
- The sources provided do not verify the number of studies or the literature coverage date of the systematic review mentioned in the proposal.
- The primary studies that would support the approximate 6% to 40% estimates for MMLU, GSM8K, HumanEval, and HellaSwag could not be checked, nor could the conditions under which those estimates might be compared.
- The available sources do not confirm the contents of a specific Contamination Transparency Card or that it has been adopted as a standard.
- Results from each method depend on the corpus, model, access, and protocol; conclusions should not automatically be generalized to other benchmarks.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction