
One-sentence definition
Contenido generado que parece plausible, pero no está respaldado por los datos, el contexto o las fuentes disponibles.
Definition: an output that is not supported
In artificial intelligence, a hallucination is an output that presents information that is incorrect, fabricated, or unsupported by the evidence available for the task. This is an operational definition: it describes the relationship between what the system produces, what it was asked to do, and the sources or data it was supposed to use. It does not imply that the system has a conscious experience, perceives something that is not there, or decides to deceive the person asking the question.
The term is used in different ways in research. Some work focuses on a lack of faithfulness to a source or supplied context; other work focuses on false or unverifiable claims. It is therefore useful to specify which criterion is being applied. An answer may contradict a document, lack support in the available sources, or be factually false. These are related problems, but they are not identical.
The distinction between “false” and “unsupported” matters. If an answer says that a contract contains a clause that does not appear in the document, that attribution is unsupported and can be checked against the text. By contrast, if an answer makes a claim for which no sources were provided, the context alone may not establish whether it is false: it could be true, false, or indeterminate. Insufficient evidence does not automatically prove the opposite.
In this guide, “hallucination” refers to a claim presented as reliable when it is not supported under the relevant criterion and evidence. That criterion might be correspondence with a text, the validity of a calculation, whether a function exists in a library, or consistency with authoritative sources. An evaluation should say what was checked, rather than simply labeling an entire answer true or false.
How it can arise—and which part of a system may fail
A generative system produces an answer based on its input, context, and learned patterns. Its output may sound coherent while including details that do not follow from those elements. An ambiguous question, insufficient information, or an instruction that assumes something false can make it harder to produce a well-grounded answer. These factors do not explain every case on their own, and the final text alone does not reveal a specific cause.
In an information-retrieval system, or RAG system, documents or passages are searched for and added to the model’s context. The chain can fail at more than one point: the search may not find the relevant source, may retrieve unsuitable material, or may provide biased passages. The model may then misinterpret what it retrieves, add claims that do not appear there, or attribute an idea to the wrong source. An erroneous answer therefore does not by itself prove that the generative model was the only cause.
External tools may also be involved: a calculator, search engine, database, or executed program. A result can be wrong because the tool returned unsuitable data, was used incorrectly, its output was misread, or the final answer described something the tool did not confirm. Evaluating the result may require examining the entire workflow, when it is available.
Multimodal tasks add image, audio, or other inputs to the mix. A description may attribute a detail to an image that cannot actually be distinguished, transcribe a word incorrectly, or present an interpretation as if it were an observed fact. Evaluation should rely on the relevant input and distinguish what is actually present from what the system infers.
Locate the possible source of the error
- 01Identify the claim in question and what the task required the system to answer.
- 02Review the original context and any retrieved sources.
- 03Check the output of tools or the interpretation of multimodal inputs.
- 04Separate retrieval failure from interpretation failure and from unsupported added claims.
- 05Record what could be verified and what remains indeterminate.
Distinctions that prevent imprecise diagnoses
A factual falsehood is a claim that contradicts verifiable facts. An unsupported claim is one for which sufficient evidence has not been established in the relevant material. It may turn out to be false, but it may also be true and simply unproven in the context examined. To identify a contradiction, an appropriate reference is needed. To point out a lack of support, it is enough to show that the source or document that was supposed to support the claim does not contain it—provided that this is the task’s criterion.
An invalid inference takes a step that does not follow from its premises, even if each premise is correct. An attribution error assigns a statement, conclusion, or data point to a source that does not support it. A nonexistent citation is an especially verifiable case: the passage can be searched for, its location checked, and its content compared with what is attributed to it. A citation that looks convincing does not prove that it is authentic or relevant.
Bias describes systematic patterns that may produce unequal results or favor certain representations. It is not synonymous with hallucination: an answer can be biased without inventing a specific fact, and an invented claim does not by itself demonstrate a pattern of bias. Randomness, meanwhile, refers to variation in outputs. The fact that two answers differ is not enough to establish that one is a hallucination; their content and evidence must be evaluated.
Disinformation generally involves spreading false or misleading information and, in some uses of the term, an intention to deceive. Calling a model’s output a “hallucination” does not prove intent. The term describes an observed result, not a human purpose on the part of the system. Similarly, disagreement between sources does not automatically show that an answer was invented: it may reflect different definitions, different dates, or a genuine controversy.
A retrieval failure is not identical to a generation hallucination either. If a system finds irrelevant documents and answers faithfully based on those documents, the problem may lie in retrieval or the corpus, even if the answer is unsuitable for the question. If the model adds information that is absent, there is also a problem of support or faithfulness. In real-world scenarios, both failures can occur together.
What is being evaluated?
| Concept | Useful question | Relevant check |
|---|---|---|
| Factual falsehood | Does the claim contradict a verifiable fact? | Compare it with reliable sources appropriate to the subject. |
| Lack of support | Does the available evidence support this claim? | Find the specific support in the context, document, or source. |
| Invalid inference | Does the conclusion follow from the premises? | Review the reasoning steps and their assumptions. |
| Attribution error | Does the cited source actually express this idea? | Locate the passage and compare its content and scope. |
| Retrieval failure | Were relevant sources found for the task? | Inspect the results, coverage, and document selection. |
Three practical examples
The following cases are hypothetical illustrations, not documented incidents or professional recommendations. In all three, the key is to compare a specific claim with evidence appropriate to the task. All three examples are discussed below in their relevant contexts. The point is not to assume that the output is wrong, but to identify what evidence would be needed to assess it. A plausible-sounding answer is only a starting point for verification, not a substitute for it. The examples also show why the scope and completeness of the material being checked should be made explicit.
They illustrate different kinds of verification: consulting an applicable medical source, examining the text of a contract, and testing software against a particular version of a library. Each requires a different standard of evidence. In every case, the result should be reported at the level the evidence supports, without making broader claims than the check allows.
Healthcare example: a contraindication that is absent from the source
Suppose someone consults a summary of drug information and the system states that one medication is contraindicated with another. The word “contraindicated” can have potentially significant consequences: it is not enough for the sentence to sound medical or to include the correct name of a drug. Verification would need to identify the relevant product and source, locate the exact warning, and assess whether the claim’s scope matches the text.
If the material being consulted does not mention the contraindication, it is fair to say that the claim is unsupported by that material. To conclude that it is false, relevant and up-to-date clinical evidence would be needed; that conclusion cannot be drawn merely from the silence of an isolated passage. Nor should this example be turned into advice for a patient: for a real possible interaction or contraindication, consult a healthcare professional or an authoritative source of information.
The evaluation can break the answer into parts: Did it name the medications correctly? Did it attribute a warning to a specific source? Does that source contain the warning? Does the wording exaggerate a precaution or present a possibility as an absolute prohibition? A medical benchmark can help compare systems on defined tasks, but a result on a test set does not settle whether a particular answer is correct or replace review proportionate to the risk.
Contracts example: attributing a clause that is not there
Imagine a tool that summarizes a contract and says it contains an automatic-renewal clause with a specific term. The most direct way to assess the claim is to search the document for the supporting text and examine the relevant sections. If the clause is absent, the system should not present it as part of the contract. The more precise conclusion may be “not found in the document provided,” which is safer than asserting that the clause does not exist in any other appendix or version.
There may also be a debatable interpretation without an invented clause. For example, an ambiguous sentence may allow more than one legal reading. In that case, the literal text, the system’s interpretation, and any legal conclusion should be distinguished. Citing a section number does not prove that the section exists or supports the summary: the reference and its context still need to be checked.
The example shows why the scope of the source matters. If only one page was supplied, the absence of a clause from that page does not establish that it is missing from the entire contract. If the complete document was supplied, appendices or external documents not included may still exist. A rigorous evaluation says what material was reviewed and avoids turning an incomplete search into a universal conclusion.
Programming example: explaining a nonexistent function
Suppose an assistant describes a function in a library and provides a call that looks reasonable, but the function does not exist in the version used by the project. A plausible name and plausible arguments are not enough to make it valid. The relevant version’s documentation can be checked, the available code inspected, or a minimal test run in a controlled environment.
Version and context are essential. A function might exist in a newer version, an extension, or a different module. Finding that it does not appear in the documentation for a particular version narrows the claim; it does not prove that the function has never existed. It is also important to distinguish a false explanation of an API from an environment error, a missing dependency, or incorrect configuration.
If the answer attributes the method to particular documentation, the citation can be checked separately. If it names no source, the code can still be tested, but a test result provides evidence only for the conditions that were tried. A snippet that works for one case is not guaranteed to be correct, safe, or compatible in every project.
How to assess a possible hallucination
There is no single check that resolves every type of task. Evaluation starts by breaking the answer into verifiable claims. A lengthy answer may mix correct facts, reasonable inferences, unsupported details, and incorrect citations. Judging it as one block makes it harder to identify what needs correction and what evidence is missing.
Next, define the standard of support: Was the answer supposed to stick to a document, provide current information, perform a calculation, describe an image, or summarize retrieved results? Each purpose calls for a different test. For a document-based question, comparing passages may be appropriate; for code, an executable test provides evidence about the behavior tested; for current facts, it may be necessary to consult up-to-date and authoritative sources.
In RAG systems, it is useful to check whether each claim is supported by the retrieved passages and whether those passages are relevant and sufficient. A provenance checker can help link claims to passages, but that link does not by itself prove that a passage is true, current, or the most appropriate source. It may also fail to resolve implicit claims, complex reasoning, or information that is not in the corpus.
Citations and references should be validated as specific objects: check that they exist, belong to the stated source, and support the claim to the extent attributed to them. An authentic citation may be irrelevant, outdated, or interpreted out of context. Conversely, a claim may be correct even when its accompanying citation is misattributed; these are separate failures.
Human review is especially important when an error could affect health, rights, finances, safety, or decisions that are difficult to reverse. The intensity of review should be proportionate to the impact and to how readily the result can be checked. A direct test may be sufficient for a low-risk task; for a high-impact decision, an automated answer should not be treated as independent verification.
A short checklist
- 01Extract the specific claims; do not assess only the overall tone.
- 02Define what evidence was expected and the scope of the task.
- 03Locate the original support and check its relevance, completeness, and date where applicable.
- 04Verify facts, inferences, citations, calculations, and tool results separately.
- 05Classify each item as supported, contradicted, or undetermined.
- 06State the limits of the review and scale up verification according to potential impact.
Mitigations: useful, but not guarantees
Retrieving sources can make evidence more readily available for answering and make claims easier to check. However, retrieval does not ensure that the right source was found or that the model uses it faithfully. Documents may be incomplete, irrelevant, or inconsistent; the system may select them poorly, interpret them incorrectly, or add content that does not appear in them.
Asking a model to say that it does not know, or to abstain when it is not confident, may reduce some speculative answers. But abstention is not an infallible detector. The system may abstain when evidence is available or answer confidently when it is not. The instruction is a behavioral measure, not an independent test of truth.
Automated verifiers can flag inconsistencies, inspect citations, or compare an answer with a set of sources. Their results depend on the task, data, and method. A verifier may miss a subtle error, share limitations with the generator, or accept an unsuitable source. Evaluations and benchmarks can measure behavior under defined conditions; they do not certify every future answer in other contexts.
A research publication from a model provider proposes that certain training and evaluation incentives may favor guessed answers over acknowledging uncertainty. This is an explanatory hypothesis put forward by that provider, not a universal explanation or an independently established consensus for every case. A proposed explanation of possible causes should be kept separate from the concrete evidence used to assess a particular output.
Related concepts and practical criteria
Hallucination is related to other concepts in the glossary, but it does not replace them. The entry on grounding concerns how an answer is supported by identifiable information or context; RAG describes an architecture that retrieves material to incorporate into generation. Neither term implies that the claims are necessarily correct.
Uncertainty concerns the limits of what can be determined or the confidence with which a conclusion is held. Expressing it clearly can help, but a statement of doubt does not guarantee that an answer is well calibrated. Fact-checking is the process of comparing claims with evidence; it can uncover errors, although its quality depends on source selection and the question being evaluated. The hallucination entry in the glossary index provides a way to find these related entries and distinguish their approaches.
As a practical rule, before accepting an answer, ask: What specific claim do I need to use? What source or test could confirm it? Is the source relevant to this case and date? Does the citation actually support what is being claimed? Which parts remain undetermined? If these questions cannot be answered, it is more precise to say that the answer has not been verified than to assign it a conclusive label.
To document a review, record the claim, the material examined, the checking method, and the result. Use bounded categories: supported by the source reviewed, contradicted by it, unsupported in that material, or not assessable from the available evidence. State the scope—for example, “in the document provided” or “in the version tested.” This wording makes it possible to correct errors without claiming more than the review establishes.
The final rule is simple: fluency, a confident tone, detail, and a formally styled citation are not evidence. Reasonable confidence comes from a check appropriate to the task, using relevant sources and making the limits explicit.
Criteria for deciding what to do
| Situation | Practical action |
|---|---|
| The claim has a specific source | Check that the source exists, is relevant, and supports the exact scope of the claim. |
| The answer summarizes a document | Compare the claims with the complete document available and note its limits. |
| The answer depends on a tool or version | Record the tool, data, and version; repeat an appropriate test. |
| The available evidence is insufficient | Mark the conclusion as undetermined and seek suitable additional sources. |
| A possible error could cause significant harm | Request expert review and do not use the automated answer as the sole basis for action. |