Chain of Thought in AI: What It Is—and What It Doesn’t Prove
01

One-sentence definition

Secuencia intermedia de razonamiento, visible o interna, que descompone un problema antes de producir una respuesta.

02

Definition: a sequence of intermediate steps expressed in text

A chain of thought (CoT) is a sequence of intermediate steps expressed in language that a model generates or receives as an example to work on a task before producing a final answer. In common use, those steps present some of the work needed to reach a conclusion—for example, separating out the facts, establishing a relationship between them, and performing an operation.

The term describes observable text and the associated prompting technique; it does not guarantee that the text is a complete or faithful transcript of the model’s internal computations. A chain can be useful as scaffolding for solving a task or as a readable explanation for a person. But the fact that it looks orderly does not show that every step is correct, that it caused the final answer, or that it reveals what happened inside the model.

Chain of thought is used especially when a task involves several steps and the model is to be guided to show them, or when examples are provided to demonstrate how to present a solution. This distinction matters to teams evaluating models: a step-by-step answer can be an object of inspection, but it does not replace checking the answer or testing the explanation’s faithfulness.

03

How it is obtained: examples, instructions, and supervision

There are several related but non-equivalent ways to obtain intermediate steps. In few-shot prompting, the prompt includes one or more examples consisting of an input, reasoning expressed in text, and an answer. The work by Wei and colleagues studies this setup and uses chain of thought to refer to the intermediate reasoning steps included in the examples.

In a zero-shot setup, the model can be asked directly to work through a problem step by step, without being shown complete chain-of-thought examples. That instruction is not the same as few-shot prompting: the information supplied in the prompt differs, even though both approaches seek to elicit intermediate steps.

A model can also be trained on data containing steps, rather than merely asked to provide them at inference time. Process supervision provides signals about intermediate steps, whereas outcome supervision focuses on the final answer. These are training and evaluation choices, distinct from adding examples or instructions to a prompt. They should not be described as though a single prompt instruction were equivalent to training a model on reasoning traces.

In all these cases, the result may be a written sequence that appears to describe the work performed. The method used to produce it does not remove the need to check it: examples can guide the format, and supervision can evaluate steps, but neither fact automatically turns generated text into a direct readout of the model’s internal state.

Three approaches worth distinguishing

  1. 01Examples in the prompt: show one or more solved problems with steps and ask the model to follow a similar format.
  2. 02Instruction in the prompt: request intermediate steps without necessarily providing solved examples.
  3. 03Supervision during training: use data or signals that evaluate intermediate steps, rather than limiting the training signal to the final result.
04

Three applied examples—and what they let you verify

The examples below are illustrative. They show how a chain can serve as a readable record while distinguishing that use from the evidence needed to claim that a result is correct or that an explanation faithfully reflects the model’s internal process.

The useful checks depend on the task: in arithmetic, an operation can be recalculated; when combining documents, the passages supporting each claim can be traced; and in debugging, program behavior can be tested. None of those checks, by itself, shows that the chain is a faithful transcript of the model’s internal computation.

05

What it is useful for—and what not to infer

A chain can act as scaffolding: it divides a task into parts that can be handled one at a time. It can also make assumptions explicit that might otherwise remain hidden in a short answer. For someone reviewing the result, that structure makes it easier to locate an operation, premise, or claim that needs checking.

However, a chain should not be mistaken for proof. It may contain an incorrect step and still arrive at the correct answer by chance, or present a plausible conclusion supported by a mistaken premise. It may also omit relevant information. The appearance of detail is not the same as the quality of the evidence.

Research on unfaithful explanations in chain-of-thought prompting shows that, in the tasks studied, chains may fail to reflect factors that influenced an answer and may rationalize it after the fact. This supports a specific caution: the text should not automatically be treated as a transparent window into the process that produced the answer. The experimental findings do not establish that every chain is unfaithful; they identify a risk that should be evaluated rather than dismissed.

One way to investigate faithfulness is to intervene on or perturb the chain and observe whether the answer changes, as explored in work by Lanham and colleagues. Such tests provide evidence about the relationship between the text and the answer under the conditions examined. They do not allow direct observation of every internal computation in the model, and their conclusions depend on the tasks and methods used.

What a chain can support—and what still needs verification

ObservationWhat it may supportWhat it does not establish by itself
The answer includes orderly steps.The model produced readable intermediate text.That every step is correct or that the chain is an internal transcript.
The calculation matches an independent operation.That the arithmetic result in this case is correct.That the chain caused the answer or faithfully represents the internal process.
The claims match passages in documents.That the answer is supported by those passages, if their scope is preserved.That the model internally interpreted the documents in the way described.
The corrected code passes defined tests.That the tested behavior meets those tests.That the debugging explanation is complete or faithful.
06

Common confusions and related concepts

Chain of thought is not synonymous with reasoning in models. Reasoning is the broader concept used to discuss tasks that involve relating information, making inferences, or solving problems; CoT, in this context, refers to intermediate steps expressed in text and to techniques that elicit or use them. A chain is an observable output, not a complete measure of a general capability.

It is also not the same as prompting in general. Prompting encompasses instructions and examples provided to a model. CoT is a prompting strategy when the aim is to elicit intermediate steps. And including a sequence of steps in a single prompt is not the same as chaining multiple separate model calls: in a chained workflow, the output of one stage can become the input to another.

In retrieval-augmented generation (RAG) tasks, a chain may articulate how a question relates to retrieved passages. The presence of those steps does not, by itself, validate the retrieval or the answer: check whether the sources really provide the evidence and whether the answer preserves its scope.

Finally, a text sequence is not the same as an executable trace. PAL and Program of Thoughts explore approaches that express some of the work as a program for an interpreter to run. A program may make it possible to verify a specific operation or result through execution, but that verification does not demonstrate the faithfulness of the textual reasoning. Step verifiers, in turn, evaluate steps or receive supervision about them; they should not simply be conflated with code execution.

Quick distinctions

ConceptFocusPractical question
Reasoning in modelsA general concept of capability or inference task.What reasoning task is being evaluated?
Chain of thoughtIntermediate steps expressed in text.What steps does the model show, and which have been checked?
PromptingInstructions and examples given to the model.What information did the model receive before answering?
RAGGeneration supported by retrieved information.Do the retrieved passages support each claim?
Executable programInstructions an interpreter can run.What behavior or calculation do the tests verify?
07

How to evaluate it rigorously

Evaluation should separate at least three questions: Is the result correct? Are the displayed steps valid? Is there evidence that those steps faithfully reflect the process that produced the answer? The same test rarely answers all three on its own.

For the result, use an appropriate independent method for the domain: recalculate an operation, check claims against documents, or run tests on the code. For the steps, explicitly review premises, inferences, calculations, and references to evidence. In tasks with multiple valid solutions, the criteria should allow correct alternatives rather than rewarding explanations that are merely plausible.

To study faithfulness, interventions or perturbations can provide information about whether changing a chain affects the answer, and under what conditions. They should be designed with controls and a clearly defined question; observing a change is not enough to conclude that the complete internal process has been recovered. Interpret results only within the limits of what the test measures.

A practical approach is to record the final answer, the visible chain, the external checks performed, and the remaining uncertainties separately. If a tool verifies a calculation, say that it verified that calculation; do not turn that into a claim that the entire reasoning process was faithful. If the documents do not support a claim, flag the gap even when the step-by-step text seems convincing.

Checklist for an evaluation

  1. 01Define what is being evaluated: the result, the validity of the steps, the faithfulness of the explanation, or an explicitly stated combination.
  2. 02Choose an appropriate independent verifier: a calculation, document cross-check, code tests, or another task-specific criterion.
  3. 03Check each relevant claim or step, and record errors, omissions, and ambiguities.
  4. 04If studying faithfulness, use interventions or perturbations with controls, and limit the conclusion to the design and tasks evaluated.
  5. 05Report the result, the checks performed, and what cannot yet be concluded separately.
08

Technical reading and limits of the evidence

The work by Wei and colleagues is a primary reference for the definition and few-shot prompting with chain-of-thought examples. Work on zero-shot reasoners studies an instruction that requests steps without being equivalent to providing solved examples. Both help distinguish prompting setups; neither, by itself, proves the faithfulness of every generated chain.

PAL and Program of Thoughts explore ways to separate the production of steps or programs from the execution of calculations. They are relevant to understanding what an interpreter can verify, especially in computational tasks; checking an executed result is not a test of internal transparency. Work on step-by-step verification studies process supervision against outcome supervision, a distinction relevant to evaluation and training.

Studies of unfaithful explanations and of faithfulness measurement examine specific risks and experimental methods. Their findings help ground caution and inform evaluation design; they do not establish a universal property of all tasks, models, or chains. The available evidence does not allow direct observation of a model’s complete internal process, so claims about faithfulness should be made within clearly stated limits.

09

Practical guidelines

When reading or requesting a chain of thought, treat it as a textual explanation that may help you inspect an answer, not as automatic proof of its correctness or internal origin. Ask for steps when they add value to the task, independently verify whatever can be checked, and retain the sources or tests that support the conclusions.

If correctness is what matters, use an appropriate independent verifier. If faithfulness is what matters, design an evaluation that intervenes on the chain and state what it does and does not measure. If all you have observed is a coherent explanation, describe it as coherent or useful for inspection: do not call it faithful without specific evidence.

This entry differs from the entry on reasoning in models, which addresses the broader concept; the focus here is intermediate text and its limits. To explore the topic further, see the related entries on prompting, AI evaluation, RAG, chain of thought, and reasoning.

10

Quick examples

11

Related concepts

12

Sources consulted