Ilustración editorial para AIME 2024: qué mide un porcentaje de aciertos en 30 problemas y por qué dos resultados aparentemente iguales pueden no ser comparables
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What “AIME 2024” exactly means

AIME 2024 is often presented as a single mathematical-reasoning benchmark, but the name compresses several decisions. In one documented evaluation implementation, the dataset combines the problems from that year’s AIME I and AIME II: thirty items in total. Each problem expects an integer answer written with three digits, from 000 through 999. Scoring can be automated because it does not depend on a human assessment of the writing or on a rubric for intermediate steps.

That composition matters from the outset. A result over all thirty problems does not rest on the same basis as one calculated over just one fifteen-problem examination. Nor should one assume that two datasets labeled “AIME 2024” use identical problem statements, identifiers, formatting adaptations, or answer keys. Evaluation harnesses can package tasks and extract answers in different ways. The benchmark label identifies a family of items; it does not, by itself, certify a shared protocol.

A rigorous reading therefore starts by turning the headline into a specific question: were AIME I, AIME II, or both evaluated; how many problems were included; and which implementation converted the model’s textual output into a scoreable answer? Without those answers, the percentage is an incomplete observation rather than a settled comparison.

02

What a score measures—and what it leaves out

AIME 2024 provides a useful signal about solving competition-style mathematics problems whose final answers can be checked exactly. The format reduces a common source of ambiguity: if the evaluator recovers the correct integer under the defined rule, the item counts as correct; otherwise, it counts as incorrect. This property makes the test attractive for running many evaluations consistently.

However, the final score does not directly observe every part of the process that led to the answer. A model may produce correct reasoning but finish in a format that the extractor does not recognize. It may also arrive at the right integer through a flawed chain of steps that exact-answer scoring does not inspect. By design, the benchmark scores the final answer recovered by the harness, not the quality of a mathematical proof.

That does not mean a strong result lacks value. It indicates performance on a bounded task involving competition-style problems and verifiable integer answers. What it cannot establish on its own is general reasoning ability, the skill to write reviewable proofs, accuracy in professional tasks, reliability in domains with incomplete information, or performance on new problem distributions. Those conclusions require additional evaluations and a demonstrated relationship between the test and the intended use case.

It is also important to distinguish mathematical competence from product usefulness. Depending on the context, a system that is useful to analysts, scientists, or engineers must interpret requirements, state assumptions, use external data traceably, detect uncertainty, and communicate limitations. None of those properties is fully measured by thirty integer answers.

What can and cannot be inferred from AIME 2024

ObservationReasonable inferenceInference not justified without additional evidence
High score under a documented protocolThe system solved many items in that dataset under those conditionsThat it reasons correctly in every domain or task
Exact final answerThe extracted integer matched the item’s answer keyThat the explanation or proof was valid
Improvement with toolsThe tools and workflow contributed performance in that runThat the base model, without that workflow, has the same level
Result over fifteen itemsThere is an estimate for that subsetThat it represents all thirty items or another examination identically
03

The minimum record that should accompany a result

A publishable claim should let another person reconstruct what was measured, even if they cannot rerun every computation. The first line of the record is the dataset: state AIME I, AIME II, or both; the number of problems; and the source or version of the problem statements. The second is system identity: the exact model name, snapshot or date, provider, and, where relevant, the reasoning configuration.

The generation mode must then be disclosed. This includes the complete prompt, or a sufficiently detailed description of its output-format instructions; temperature and other sampling parameters; the maximum token limit; time limits; the number of calls per problem; and any interruption or retry policy. When these details are absent, a reader cannot distinguish a one-sample result from an intensive search through many trajectories.

The record must also separate the model from its environment. Was a calculator, code execution, web browsing, document retrieval, or another tool available? Could the system check its own answers? Did an external procedure choose the best response? These conditions may be appropriate when they reflect intended use, but they should appear beside the percentage rather than remain hidden behind the model name.

Finally, the harness needs to be described. One documented AIME 2024 implementation specifies an output prompt in the form “ANSWER: $ANSWER,” exact-match scoring, and scorer changes related to LaTeX and empty responses. Details such as these show why the parser is not an administrative matter: it determines which text becomes an integer and which output is declared invalid. The policy for missing answers, multiple answers, ambiguous text, and formatting errors should be explicit.

Process for turning a headline into an auditable record

  1. 01Identify the exact dataset and count the items that were actually scored.
  2. 02Record the model version, evaluation date, and generation parameters.
  3. 03Classify the result as one sample, multiple samples, consensus, reranking, or a combination.
  4. 04Note tools, token budget, time, and total number of calls.
  5. 05Document the prompt, extractor, error handling, and aggregation formula.
  6. 06Label comparability with other results as direct, partial, or not established.
04

One evaluation can produce many different figures

Pass@1 answers a simple question: with one sample per problem, how many final answers were correct? It is usually the interpretation closest to a single interaction, provided that the prompt, temperature, budget, and tools are also specified. Even then, it is not necessarily a portrait of an ordinary conversation: it may include instructions designed specifically for the benchmark.

Pass@k changes the question. Rather than one opportunity, several answers are generated for each problem, and the measure asks whether any one of them is correct. This can be useful for studying the system’s search potential, but it rises with the number of attempts and is not equivalent to the likelihood that one answer delivered to a user will succeed. Reporting pass@k without reporting k makes the figure impossible to interpret.

Majority voting or consensus generates multiple answers and selects the one receiving the most support under a defined rule. It can improve stability when several trajectories reach the same solution, but it does not guarantee truth: attempts can share the same mistake. Its performance depends on the number of samples, how equivalent answers are grouped, and what happens to formats that cannot be extracted.

Reranking adds another component: after candidates are generated, a selector chooses one. That selector may be the model itself, a verifier, a heuristic, a tool, or a different system. A reranked result therefore measures an entire pipeline. It should not automatically be attributed to the isolated generating model. An OpenAI publication explicitly distinguishes one-sample results, consensus across dozens of samples, and reranking across a much larger number of samples; that distinction is methodologically essential.

Tools introduce another branch in the protocol. A calculator or code environment may reduce arithmetic errors; external search may supply information not contained in the prompt. Neither option is inherently illegitimate, but “with tools” and “without tools” answer different questions. Model cards that report both conditions provide a more informative reference than a single figure.

05

Why an apparently identical percentage may not represent the same number of correct answers

In a thirty-problem set, every correct answer accounts for a meaningful share of the total. If a percentage is reported with decimals, it may come from a single run, repeated runs, an average across subsets, or aggregation across configurations. In a fifteen-problem set, the jumps in a single execution are larger still. A value such as 80% therefore does not automatically reveal an integer number of correct answers or the number of problems evaluated.

Rounding is only part of the issue. One provider may report the average of many runs; another may report the best outcome from one run; a third may average per-question results after sampling multiple times. Those choices can be defensible for different purposes, but they need to be named. A figure without a denominator, dispersion, or aggregation procedure should not be treated as a precise measurement.

Comparability requires holding constant—or at least disclosing—the elements that affect effective difficulty: dataset, problem-statement version, model, prompt, tools, generation, extraction, and scoring. If any of these differ, the comparison may still be indicative, but it should not simply be converted into a capability ranking.

A practical comparability rule

SituationVerdictHow to communicate it
The same thirty items, the same harness, one sample, and the same tool availabilityRelatively directly comparableState model versions and date
The same dataset, but pass@1 versus consensus or rerankingNot comparable as one-sample capabilityCompare only as pipelines, including cost and sample count
Fifteen items versus thirty, even though both are called AIME 2024Partially comparableShow denominators and avoid a single ranking
The same percentage without the prompt, parser, or error policyNot establishedRequest documentation before drawing conclusions
With tools versus without toolsNot comparable as an isolated modelUse separate columns and describe the tools
06

Prior exposure, public availability, and saturation

AIME 2024 should also be examined as a dataset that has been publicly available. Work on evaluating uncontaminated mathematics competitions treats that availability as a reasonable source of concern: a model or retrieval system may have been exposed to problems, solutions, discussions, or variants during development. This does not demonstrate that any particular model was trained on those items, nor does it allow every strong result to be attributed to memorization.

The cautious formulation is conditional. Prior exposure is a validity risk that should be disclosed when sufficient information about training data, fine-tuning data, retrieval, and cutoff dates is unavailable. The absence of public evidence of exposure does not prove independence either. Between those poles lies uncertainty, not an automatic conclusion.

A provider can be asked for proportionate evidence: the training-data cutoff date, a description of evaluation-data filtering, its policy on competition materials, the model’s availability date, and details of any retrieval tools. If it cannot provide this information, the result can remain contextual evidence, but with a clear limit: it should not function as the sole proof of generalization to unseen problems.

Saturation has a further practical consequence. The more popular a benchmark becomes, the more likely it is that prompts, solutions, strategies, and evaluation configurations circulate widely. For product decisions, treat AIME 2024 as a historical signal and complement it with a private, recent battery that represents the real work, while respecting the competition materials’ policies on use and reproduction.

07

How to use AIME 2024 in a product decision

For a technical decision-maker, AIME 2024 can serve as a secondary signal when shortlisting systems that solve structured mathematics problems. It is most useful when paired with a clear protocol and when compared with results produced under the same conditions. It should not become a single threshold for procurement, deployment, or safety.

The complementary evaluation depends on the use case. If a product generates quantitative analyses, relevant tasks should use the organization’s own data, units, assumptions, and calculation review. If it must explain results, assess clarity, traceability, and error detection—not merely the final integer. If it operates with tools, measure the complete pipeline, including permissions, cost, latency, tool failures, and verification. If the concern is expert knowledge, a benchmark such as GPQA Diamond can provide a different signal, but it does not replace tests of the specific workflow.

It is equally important to separate efficiency from outcome. Two systems with similar accuracy may require very different numbers of samples, tokens, calls, or time. In production, those differences affect cost, latency, throughput, and predictability. A benchmark table without an inference budget can hide a difference that is decisive for the use case.

Readers comparing a DeepSeek-R1 record, a GPT-6 Astra record, or other models should apply the same discipline and should not infer unpublished conditions from the brand. A commercial name does not replace a snapshot, and a figure attributed to an organization does not replace protocol documentation. The benchmarks index and the dedicated AIME 2024 record are the appropriate places to retain these conditions alongside every result.

Final checklist before reusing an AIME 2024 figure

  1. 01Is it known whether the result covers fifteen or thirty problems, and which ones?
  2. 02Does it identify the exact model or snapshot and execution date?
  3. 03Is it pass@1, pass@k, consensus, reranking, or a mixed method?
  4. 04How many samples, tokens, calls, and how much time were used per problem?
  5. 05Were tools or external verifiers available?
  6. 06Are the prompt, harness, parser, and policy for unparseable outputs published?
  7. 07Does the percentage come from one run, an average, or the best observed value?
  8. 08Was uncertainty regarding prior exposure and public availability disclosed?
  9. 09Is the decision also supported by representative, recent, first-party evaluations?

Open questions

  • Available sources do not establish whether a specific model was trained, fine-tuned, or previously evaluated on AIME 2024 problems or solutions.
  • Not all public results disclose the prompt, temperature, token budget, parser, number of repetitions, or handling of unparseable responses.
  • An isolated figure attributed to a model does not reveal the cost, latency, or reliability of a production pipeline.
  • Comparability across different AIME 2024 implementations may be partial or not established even when they share the same benchmark name.
08

Keep exploring

08

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction