Ilustración editorial para Cómo diseñar una evaluación propia de IA: del caso de uso a un umbral de despliegue verificable
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

Public benchmarks are an initial signal, not a deployment decision

At best, a benchmark table answers a narrow question: how a model performed under a particular protocol, with a particular task set, version, and scoring rules. It does not, by itself, answer whether an assistant will be useful, safe, fast enough, or affordable within a real process. The distinction matters because a product is not made up of a model alone. It includes instructions, tools, document retrieval, output formats, access controls, an interface, human supervisors, and exception procedures.

Benchmarks can help narrow the candidate set and formulate hypotheses. For example, a public coding-oriented test may be relevant when the work resembles resolving software issues; a terminal test may provide useful information when an agent must operate in terminals and real environments; and an interface-focused test may be closer when the task requires using graphical interfaces. However, apparent similarity does not remove the need to test your own workflow with the constraints and data the system will actually encounter.

Practical literature on model benchmarks warns of limitations such as data contamination, configuration differences, and results that do not capture latency or tool integration. Those limitations do not invalidate public tests: they define the kinds of inference they support. The recommended approach is to use their results as external context and reserve the product decision for a local, repeatable, documented evaluation.

To explore that distinction, the editorial cluster should include the resource “What public benchmarks can—and cannot—contribute.” This guide should also link from the learning hub page, in the “Guides for implementation and verification” block, and receive a discovery link through the module “Before choosing a model: evaluate it with your use case.” Publication should depend on those destinations and their navigation being implemented or planned within the same cluster.

02

Translate the problem into an evaluable hypothesis

The starting point should not be “which model gets the highest score,” but a verifiable description of the task. Identify who uses the system, what input they provide, what outcome they need, what subsequent actions depend on that outcome, and what happens when the system fails. An incorrect classification of a commercial inquiry may require correction; an answer that falsely attributes a claim to an internal document may lead to a wrong decision; an unauthorized external action can have more serious consequences. The metric and threshold should reflect that difference.

A useful hypothesis takes a conditional form: for a defined segment of users and tasks, the system produces an output that meets explicit criteria at a quality, time, and cost level compatible with the process, without exceeding an error budget. This statement forces a decision about what “meets” means. In structured extraction, it may mean valid and correct fields. In a summary, it may require coverage of relevant points, no fabricated content, and attribution to the available sources. In an agent, it may mean completing a task without violating permissions or requiring more human intervention than anticipated.

Microsoft describes a framework for prioritizing use cases that combines business viability, experience and desirability, and technical feasibility. It is a prioritization framework, not a test of model quality; nevertheless, it helps avoid a frequent omission: assessing technical capability without confirming that the use case has a plausible process, owner, data, user experience, and operational benefit.

Minimum hypothesis brief

ElementQuestion it must answerExpected evidence
User and contextWho uses the result, and at what point?Current workflow and case segmentation
TaskWhat input does it receive, and what output or action does it produce?Anonymized examples and output contract
Acceptable outcomeWhat must be correct, and what can be delegated to review?Rubric and acceptance criteria
Tolerable harmWhich error blocks use?Risk and escalation record
DecisionWhich threshold allows deployment, iteration, or rejection?Predefined results matrix
03

Define and version the complete system

Reproducibility requires stating what was compared. Record the provider and the model version or identifier where available, the execution date, generation parameters, the instruction template, system message, available tools and their versions, output schema, retrieval index, document version, and permission rules. Add retry logic, validation, fallback, and human escalation. If this configuration cannot be reconstructed, a difference in results between two runs will be difficult to interpret.

Not all components change at the same rate. A provider may update a service, internal documents may change daily, and a team may change an instruction to address a failure. The record does not prevent those changes, but it makes it possible to know what changed before attributing an improvement or regression to the model. It is especially important not to compare one option using updated retrieval with another using an older index, or to attribute to a model variant a result caused by a different prompt.

Freezing does not mean making the product static. During an evaluation round, it means fixing a candidate configuration and a test set before observing results. Once the round is complete, a new version can be proposed, but it should be evaluated again against the same reference or against an explicitly updated reference. This discipline reduces the risk of retrospectively choosing the prompt that happened to perform well on the test.

Versioning process for every run

  1. 01Assign an identifier to the run and fix the date, environment, and owner.
  2. 02Capture the model, parameters, prompts, tools, permissions, schema, retriever, and document index.
  3. 03Run the set without changing cases or criteria during the round.
  4. 04Store outputs, permitted traces, errors, costs, and latencies with references to the identifier.
  5. 05Record every exclusion, retry, or human intervention and its reason.
  6. 06Approve, iterate, or discard through a decision linked to that evidence.
04

Build a representative, separated evaluation set

An evaluation set should represent operational decisions and conditions, not only easy or memorable examples. Start by sampling historical tasks while accounting for relevant segments: language, length, document type, input channel, degree of ambiguity, common cases, and high-impact cases. Then add deliberately difficult cases: contradictory instructions, incomplete information, outdated documents, malformed inputs, out-of-scope requests, and situations where the correct outcome is to abstain or ask for clarification.

Separate at least two partitions: development and test. The development partition is used to design prompts, rules, tools, and rubrics. The test partition is held back for a final comparison or defined milestones. The methodological reason is straightforward: if the system is repeatedly adjusted after seeing results on the same cases, those observations stop measuring generalization and become part of development. A third validation partition can be useful for projects with sufficient data, but it is not a universal requirement; the important point is to document how each case is used.

Private data adds further practical obligations. Minimize the personal data and secrets included in the evaluation, apply access controls, and avoid moving sensitive content into unauthorized environments. This guide does not itself determine whether a processing activity is lawful or what providers’ contractual requirements are. Before incorporating private documents, use the planned resource “Example evaluation matrix for sensitive documents” and validate applicable conditions with legal, security, and data protection functions.

Synthetic cases can expand coverage when examples are scarce, but they do not automatically substitute for operational reality. They should be labeled as synthetic, reviewed, and kept separate in reports. If only cases created for a demonstration work well, the evaluation does not provide sufficient evidence about production behavior.

05

Choose metrics that match the task and the risk

There is no single metric for generative systems. Exact-match scoring can be reasonable when the output is a closed label or exact value, but it is often insufficient for summaries, open-ended answers, or action plans. In those cases, a human rubric can assess separate aspects: factual correctness, coverage, clarity, instruction compliance, appropriate use of sources, and acknowledgment of uncertainty. Automated evaluation can complement that work through schema validators, business rules, required-field checks, and safety tests; it should not conceal which properties it cannot verify.

For retrieval-augmented generation with sources, measure retrieval and generation separately. Among other signals, assess whether relevant evidence was retrieved, whether important claims are supported by the available evidence, whether citations or references match the content, and whether the system abstains when evidence is insufficient. A fluent answer does not demonstrate faithfulness. The planned resource “How to evaluate faithfulness, coverage, and citations in RAG systems” can examine these criteria in more depth.

For structured outputs, measure schema validity, field presence, the semantic accuracy of each field, and recovery from an error. Formally valid JSON can still classify incorrectly; a correct classification may still be unusable if a required field is missing. The planned resource “How to measure schema validity and error recovery” is intended to support the design of these tests.

Add operational metrics: latency by percentile and segment, cost per completed task, retry frequency, fallback rate, human intervention, and end-to-end success. Report distributions as well as averages where possible, because an average can hide latency tails or segments with consistently worse outcomes. Thresholds cannot be derived from a universal figure: they depend on the process, volume, harm, and supervision capacity.

Metric decision matrix

Outcome typePrimary measuresComplementary check
Closed label or extractionPer-field accuracy; precision and coverage where appropriateSchema validity and business rules
Summary or open-ended responseHuman rubric for correctness, coverage, and clarityReview of unsupported claims
RAG with sourcesRetrieval relevance; faithfulness and coverageCorrespondence between claim and source
Agent with toolsTask success; unnecessary steps; interventionsPermission violations, confirmations, and reversibility
OperationsLatency, cost, retries, and fallbacksResults by segment and critical cases
06

Design the rubric and manage human disagreement

A rubric turns broad judgments into observable decisions. Rather than asking “is the response good?”, define dimensions, scales, and examples. For the feedback assistant, a faithfulness dimension could distinguish between: every relevant claim supported by the text; minor acceptable inferences that are identified; unsupported claims; and clearly false attributions. Another dimension can assess whether category and priority follow the team’s operational definitions.

Some controls can be automated deterministically: schema parsing, length limits, prohibited fields, matching against a permission list, or the presence of required references. Others require interpretation. An evaluator based on another model can speed up screening, but it should be calibrated against human annotations on a representative sample and its errors should be analyzed. It should not be treated as an independent arbiter or used to replace human review where consequences are high.

To measure agreement, two evaluators can be compared using percentage agreement on simple criteria or a measure of agreement that accounts for agreement expected by chance, such as kappa, provided the scale and distribution allow it. There is no universal level of disagreement that invalidates an evaluation. As an operating rule, review the rubric if disagreement affects cases that would determine deployment, if it is concentrated in a critical dimension, or if evaluators interpret the criterion incompatibly. Document the revision and, if the rules change, reevaluate the affected cases.

07

Compare fairly and decide using explicit thresholds

Define thresholds before running the final comparison. Establish absolute blockers, minimums by segment, and operational objectives. A blocker may be an output that exposes unauthorized information, an action performed without required confirmation, a permission violation, or a critical unsupported claim in a context where the system is presented as source-grounded. A minimum by segment prevents an acceptable aggregate result from concealing unacceptable behavior for a language, document type, or user group.

A decision can have three outcomes: deploy with controls, iterate, or discard. Deploying with controls is appropriate only if blockers are zero within the evaluated coverage, agreed minimums are met, and supervision is compatible with residual risk. Iteration requires identifying what will change and how it will be measured again. Discarding may be the responsible conclusion if no configuration reaches the threshold within cost, latency, or safety limits.

For agents that perform external actions, the evaluation must explicitly test permissions, confirmations, scope, and reversibility. It is not enough for the task to end successfully. The planned resource “Testing permissions, confirmations, and reversibility before giving an agent actions” should cover this type of validation. This guide does not set specific regulatory requirements under the European AI Act: the verified sources provided are not sufficient primary legal text to determine applicable articles, dates, or obligations. That review should be conducted using current legal sources and specialist advice.

Deployment gate

  1. 01Confirm that the configuration, set, and rubric are frozen and identified.
  2. 02Run every case, including safety and abstention cases.
  3. 03Apply blockers first, then minimums by segment.
  4. 04Manually review critical errors and evaluation disagreements.
  5. 05Compare quality, latency, cost, and human intervention against the target process.
  6. 06Record the decision, owner, limitations, and mandatory reevaluation date.
08

Monitor after deployment and retain the decision record

Deployment does not turn evaluation into a closed administrative step. User inputs, retrieved documents, provider models, tools, and abuse patterns change. Instrument traces with appropriate data minimization so an output can be linked to its configuration, retrieval, tools, response time, cost, validations, fallback, and human intervention. The planned resource “How to instrument traces, latency, and cost per task” can serve as the operational continuation.

Define warning signals: an increase in schema errors, a drop in success by segment, growth in abstentions or retries, delays in high percentiles, rising cost per task, reviewed complaints, and permission events. Sample production outputs for human evaluation using procedures that respect applicable data controls. Schedule reevaluations after material changes and periodically as well, even when no change is declared, because an external dependency may vary.

The final template should bring together an evaluation brief, the results matrix, and a decision record. The brief includes the objective, segments, risks, versioned configuration, data provenance, and metrics. The matrix shows aggregate and segment-level results together with critical cases. The record explains what evidence was reviewed, which thresholds were applied, what limitations remain, who approved the decision, and when testing must be repeated. For closing navigation, add “Next step” to a downloadable template or a companion article on observability. This closing element must not be published as an active link if its destination does not exist or is not planned within the cluster.

Open questions

  • The verified sources provided are mostly secondary guides. They support practical recommendations, but they do not establish universal thresholds or definitive regulatory obligations.
  • No current primary legal source was provided to specify European AI Act articles, dates, and applicability for the cases described.
  • The actual existence of internal destinations, the downloadable template, and the companion observability article has not been verified; their activation should depend on cluster planning.
  • The information provided does not allow confirmation of current provider terms regarding versioning, data retention, or model changes. These should be verified before publication makes specific claims about them.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction