Ilustración editorial para RAG con evidencia: cómo decidir qué responder, qué citar y cuándo abstenerse
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

Retrieving a passage does not prove an answer

In a retrieval-augmented generation (RAG) system, the result depends on more than whether search finds documents related to a query. It also depends on whether the final answer uses those documents correctly. A passage may cover the same topic without answering the question; several relevant passages may leave an important part unresolved; and a reference may be genuine without supporting the claim next to it.

So the design question should not be only, “Did the system find documents?” It should also be, “What claims does the retrieved evidence allow the system to make?” The change may seem small, but it affects how answers are constructed, how the system is tested, and how failures are investigated. The goal is not to force an answer to every query. It is to answer when the passages are sufficient and choose a controlled response when they are not.

This guide focuses on that final decision. It does not cover corpus maintenance or versioning, embedding migrations, or the evaluation of web-research agents. Nor does it suggest that an automated score can certify an answer as true. Instead, it offers an operational method for inspecting the relationship between the query, the evidence, and the claims—and for finding situations where that relationship breaks down.

02

Define an evidence contract before evaluating

An evidence contract describes what an answer must satisfy to be acceptable in a particular product. It is not a promise of absolute accuracy, nor a generic instruction such as “use the sources.” It should make clear which sources the system may use, what counts as sufficient support, which references it must return, and what it should do when the evidence falls short.

Start by classifying the questions you expect to receive. A question about a date, an internal procedure, a comparison, or a recommendation may require different kinds of support. For a simple factual question, one direct and applicable passage may be enough. A comparison may require information about both items. For a compound question, every part must be supported or explicitly marked as unresolved. This is a product decision: the available sources do not establish a universal threshold for sufficient evidence.

Specify admissibility as well. For example, an internal assistant might be limited to approved documents in its workspace; a product that answers questions about public documentation might allow a selected set of publications. Do not silently mix an admissible source with one the product should not use. If two admissible sources disagree, decide whether the system should surface the conflict, apply an approved precedence rule, or abstain. Do not let it invent a resolution.

Finally, define what a person needs in order to review a reference: a stable document identity and a findable location, such as a section or passage. The specific technical format will depend on the system. What matters is that a reference is not merely a decorative label and that a reviewer can check the cited content without having to guess where it is.

Decisions the evidence contract should settle

DecisionDesign questionPractical criterion
Admissible sourcesWhich documents may the system use to answer?Define the permitted set and how to handle material outside it.
Sufficient supportWhat evidence allows the system to answer this kind of question?Require direct support for each relevant factual part.
ConflictsWhat happens when admissible sources disagree?Surface the conflict or apply an explicit rule; do not conceal it.
ReferencesCan a person locate and review the passage?Return identifiers and locations that are useful for inspection.
Missing evidenceWhich controlled response is allowed?Choose among a partial answer, clarification, or abstention.
03

Separate claims from their references

An answer can combine several claims: a fact, a condition, a date, and a conclusion. If they all share a list of references at the end, it can be difficult to tell which source supports which claim. Design the output so a reviewer can inspect the verifiable claims and the passages associated with them separately.

There is no need to impose one visual format. The response could be prose with references beside each claim, a structured list of claims and sources, or a short answer accompanied by passages that can be inspected. Whatever the format, it should preserve the connection between each claim and its support. Avoid grouping distinct factual statements under a single citation if the passage supports only part of them.

The distinction between a correct reference and an answer that is faithful to the evidence matters. A document may be relevant and authentic, yet the answer may attribute something to it that it does not say, or rely on knowledge beyond the passage. Research on RAG attributions addresses precisely the difference between citation correctness and faithful use of evidence, including the problem of adding citations after drafting. The practical consequence is that a reviewer should compare the claim with the text, rather than checking only that the document exists.

For longer answers, consider dividing the text into units that can be reviewed. For each factual unit, record whether support is direct, partial, contradictory, or insufficient. This is a design recommendation, not a guarantee that the model will always classify evidence correctly. The value of separating claims is that it makes failures visible and lets the team address them with concrete examples.

04

Test retrieval, support, and the response decision separately

A useful evaluation does not collapse the whole process into one overall score. When an answer fails, the team needs to distinguish whether the necessary passage was not retrieved, whether it was retrieved but misinterpreted, whether the reference cannot be inspected, or whether the system should have abstained. Microsoft’s RAG evaluator documentation describes separate dimensions for examining retrieval and aspects of the response such as grounding, relevance, and completeness. RAGAS, a published RAG evaluation framework, also separates dimensions such as context relevance, answer faithfulness, and answer relevance.

You do not need to begin with a huge test suite. Assemble questions that represent real use and prepare an evidence-based expected result for each one: what part should be answerable, which passage would be sufficient, and what output is appropriate if that passage is missing. A person should review the reference material and criteria; otherwise, the test set could precisely measure the wrong interpretation.

Record results by stage. For each question, note which documents and passages were retrieved, which claims the system made, which references it attached, and which response decision it took. At a minimum, distinguish insufficient retrieval, irrelevant passages, partial evidence, unsupported claims, unfindable citations, unreported conflicts, and inappropriate abstention. These categories are a diagnostic proposal: adapt the names to your product, but preserve the ability to locate where a failure occurred.

Do not interpret an aggregate score as a certificate of truth. A favorable relevance result does not prove that every claim is supported; a well-grounded answer can still be incomplete; and a verifiable reference does not prove that the conclusion is correct. Automated measures help detect patterns and prioritize reviews, but boundary cases and high-impact errors require human inspection appropriate to the product’s risk.

Initial testing protocol

  1. 01Choose representative questions and write down the specific evidence that would allow each one to be answered.
  2. 02Manually review the expected passages and note whether they support a complete answer, a partial answer, or no answer.
  3. 03Run the system and retain the query, retrieved passages, answer, references, and response decision.
  4. 04Compare every factual claim with its associated passage; classify support as direct, partial, absent, or contradictory.
  5. 05Record the failure type and change one thing at a time: retrieval, instructions, or response logic.
  6. 06Rerun the test set after the change and review examples that previously passed as well.
05

Include difficult cases, not just answerable questions

A test set made up only of questions whose answers appear verbatim in a document does not show whether the system knows when to stop. Deliberately include queries with no answer in the corpus, ambiguous queries, questions that can only be answered in part, and passages that share vocabulary with the query but do not provide the requested information. These cases reveal whether the model mistakes topical similarity for proof.

For an unanswered question, check whether the system recognizes the limits of the available material without claiming that the fact does not exist in the world. “I cannot find sufficient support in the available documents” is different from “That does not exist.” For ambiguity, decide whether a clarifying question would resolve an important difference, such as a time period, region, or definition. For a partial answer, preserve what is supported and explicitly delimit what is not.

Contradictory evidence deserves a dedicated test. Retrieve two admissible documents that make incompatible claims and observe whether the answer presents one version as indisputable, combines them incoherently, or surfaces the discrepancy. If an approved rule exists for determining currency or precedence, check that the system applies it only when the evidence allows it. If no such rule exists, the prudent response is to describe the conflict or abstain from resolving it.

Test distractors too: documents that mention the same entities or concepts but do not answer the question. It is not enough to check whether the system retrieved something related. Review whether the claims actually follow from that content. The aim is not to penalize the system for failing to guess; it is to check whether it can distinguish a topical clue from sufficient support.

A compact adversarial test matrix

CaseWhat to checkExpected response
No answer in the corpusWhether the system avoids filling gaps with unsupported knowledge.State that the available documents are insufficient.
Ambiguous questionWhether the system identifies the missing interpretation and its importance.Ask for clarification when the interpretation changes the answer.
Partially supported answerWhether the system separates available facts from the unresolved part.Answer only the supported part and delimit the rest.
Contradictory sourcesWhether the system recognizes the discrepancy rather than hiding it.Surface the conflict or apply an explicit, verifiable rule.
Topically similar passageWhether the system distinguishes lexical similarity from direct evidence.Do not claim what the passage does not establish.
06

Check each citation against the specific claim

Citation review can follow a short sequence. First, verify that the reference points to an admissible source and that the document exists in the environment where it will be used. Second, check that the location makes the passage findable. Third, compare the passage with the exact claim: does it support the claim directly, provide context only, contradict it, or say nothing about that point? Finally, check whether the answer added conditions, causality, or scope that the source does not contain.

This review must be local. A document’s general reliability does not mean it supports every claim attributed to it. Likewise, an answer can sound reasonable and have a relevant reference without the cited passage implying its conclusion. When the source supports only part of a claim, revise the wording to reflect that limit or remove the unsupported remainder.

Test defective references too: nonexistent identifiers, overly broad locations, duplicate passages, and citations attached to the wrong claim. If the interface displays references at the end of a paragraph, a reviewer should be able to determine clearly which sentences they cover. A technically functional link does not replace that semantic relationship.

When recording a failure, preserve the exact query, passage, and answer. “The citation is wrong” is not enough to improve the system: say whether the source does not exist, the passage cannot be found, its content does not support the claim, or the answer overstates what can be concluded. This precision helps determine whether retrieval, answer composition, or attribution rules need to change.

07

Choose whether to answer, answer partially, clarify, or abstain

The response does not have to be binary. A practical policy distinguishes four options: answer when the passages support what was asked; answer partially when one part is resolved and another is not; ask for clarification when different reasonable interpretations would lead to different answers; and abstain when evidence is insufficient or a conflict cannot be resolved using the available rules.

Define these options with examples and observable criteria. Avoid vague instructions such as “answer cautiously”: they do not tell the system what to do when a date is missing, sources conflict, or a question is ambiguous. A useful abstention should explain the limitation narrowly and, if possible, say what information would allow the system to continue. It should not invent an explanation for why it did not answer.

Evaluate errors in both directions. An unsupported answer is a failure, but unnecessary abstention can prevent the user from receiving information that was available. So the test set should include both cases where abstention is appropriate and cases where the evidence is sufficient to answer. Microsoft’s generative AI evaluation documentation includes abstention among the dimensions that can be evaluated in generative systems; here, the principle is applied to the response decision without extending the scope to web-research agents.

Partial answers need particular care: do not present a confirmed part in a way that makes it seem to resolve the whole query. If someone asks about a duration and an exception, give the supported duration and state that the passage does not identify who authorizes the exception. If the unresolved part substantially changes the meaning, ask before offering a conclusion that could mislead.

Response decision rule

Evidence statusActionWhat the person should see
Direct and sufficient for the whole questionAnswerDistinct claims and findable references.
Direct for one part, insufficient for anotherAnswer partiallyWhich part is resolved and which remains open.
The intended interpretation is unclear and affects the resultAsk for clarificationThe specific ambiguity that needs to be resolved.
Absent, irrelevant, or contradictory with no resolution ruleAbstain or surface the conflictThe limits of the available evidence, without presenting them as universal certainty.
08

Analyze failures before changing the system

When a test fails, avoid immediately trying to fix it with a longer instruction. First determine which stage caused the problem. If the necessary document did not appear, retrieval may be at fault. If the passage appears but does not support the answer, interpretation, answer composition, or the rule authorizing a response may be at fault. If the reference does not lead to the passage, the problem is traceability. If evidence is insufficient and the system answers as though it were enough, the abstention decision has failed.

Keep a record of failed examples, including their category, impact, and the change applied. When possible, change one component at a time and rerun both the failing case and examples that previously passed. A change that improves unanswered cases may increase unnecessary abstentions; one that encourages fluent answers may conceal limitations. Regression review helps reveal these trade-offs.

Separate indicators by stage. You could report how many queries retrieve passages judged relevant, how many claims have sufficient support, how many references can be inspected, and how the policy behaves when evidence is insufficient. Do not combine everything into one number that hides these distinctions. None of these measures, by itself, proves the system is correct; they help the team decide what to inspect and compare changes using the same test set.

Review errors according to their effects in the intended use. An unsupported claim about an internal procedure may have different consequences from an error in a low-impact question. The level of human review, the size of the test set, and acceptance criteria should be adjusted to the product’s risk and context. This calibration is a decision for the responsible team, not a universal threshold derived from a metric.

Diagnosing a failed result

  1. 01Did the necessary passage appear among the retrieved results? If not, investigate retrieval.
  2. 02Does the passage address the specific question or merely share its topic? If it is only topical, classify it as insufficient evidence.
  3. 03Does each factual claim follow from its associated passage? If not, correct the attribution or the answer.
  4. 04Can a person locate the passage using the reference? If not, fix traceability.
  5. 05Was the evidence insufficient even though the system answered? Review the rules for partial answers, clarification, or abstention.
  6. 06Rerun the case and regression examples; keep the result and an explanation of the change.
09

Acceptance checklist before deployment

Before putting a RAG system into use, check that the team can answer the acceptance questions clearly. It is not enough for a few demonstrations to produce plausible answers. The system must be tested with answerable and unanswerable queries, retrieval quality must be distinguished from claim support, and the response decision must be checked against the available evidence.

The review should include cases that often fall outside a demonstration: conflicts between sources, incomplete questions, retrieval of a similar but insufficient document, and queries whose answers are only partially documented. Also check that the format makes it possible to inspect the source and connect it to the corresponding claim. If this review cannot be done, the team does not yet have a sufficient basis for saying that citations make the output auditable.

Citations are an aid to inspection, not a guarantee of truth or correctness. A citation can point to the right document without supporting the sentence; a passage can support a claim without resolving everything that was asked; and an automated evaluation can simplify important nuances. The operational criterion is narrower and testable: require support for every important factual claim, describe limitations when support falls short, and measure separately where the process fails.

The system is better prepared when the team can show reviewed examples of all four decisions—answer, answer partially, ask for clarification, and abstain—identify the evidence supporting each claim, and explain how contradictions were handled. If it cannot do that yet, the next step is not to add more citations to the text. It is to improve the evidence contract, the test set, or evidence traceability.

Open questions

  • The supplied sources do not establish a universal threshold for sufficient evidence; it must be defined according to question type and product risk.
  • The available sources support separating evaluation dimensions and grounding answers, but do not prescribe a single reference format.
  • The appropriate approach to resolving contradictions depends on authority or currency rules that each team must define and verify.
  • Automated metrics can help compare systems, but they do not replace human review of boundary cases or high-impact cases.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction