What tasks does it contain?
Coverage, difficulty, language, topicality and possible contamination.
SampleA metric is only useful when we know the task, the conditions, and the error that matters.
The result begins with the design of the test.
Coverage, difficulty, language, topicality and possible contamination.
SampleAccuracy, human preference, cost, time or success rate.
AimPrompt, tools, reasoning, samples and budget.
EquityTurn real needs into a small, repeatable set.
Includes normal cases, edges and costly failures.
CasesSeparates correctness, usefulness, security and format.
CriterionModel, recovery, tools, latency and cost.
ProductionA public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
A moving alias and a snapshot may produce different results. Record the model, date, and provider.
Prompt, reasoning, tools, number of attempts, and budget must be equivalent for comparison.
A small difference may disappear between runs. Keep samples, dispersion, and failures, not just the average.
Validate whether the improvement holds across languages, formats, and real cases close to your product.
Referencias utilizadas para ampliar y revisar esta página.