What Is Chain-of-Thought Prompting?

Chain-of-thought is prompting a model to produce intermediate reasoning steps before its final answer, rather than answering directly. On multi-step problems — arithmetic, logic, planning — this materially improves accuracy, and the effect appears mainly in sufficiently large models. The working explanation is that generating the steps lets the model allocate more computation to the problem and condition each step on the last, approximating a worked solution rather than a single-shot guess.

What it does
Elicits intermediate steps before the answer
Introduced
Wei et al., 2022
Caveat
The stated steps may not be the real cause

More computation, not necessarily honest introspection

The practical mechanism is that each generated step becomes context for the next, so the model spends more forward passes on a hard problem and builds its answer incrementally. This is why it helps most where a problem genuinely decomposes into steps, and little on single-fact recall.

A research-grade caveat matters here: the written chain is not guaranteed to be the model’s actual reasoning. Studies show a model can produce a plausible rationale that does not reflect what drove its answer, and can even rationalise a biased conclusion. For an agent acting on its own output, a chain-of-thought is a useful artefact to inspect, not a trustworthy audit of why it decided.

Related standards

Wei et al., 2022 — Chain-of-Thought Prompting Elicits Reasoning

Questions

Does chain-of-thought work on any model?

The benefit is strongest in larger models; small models often gain little, and the original work framed it as an emergent behaviour of scale.

Is the reasoning trace trustworthy?

Not inherently — it can be post-hoc rationalisation. It is evidence to examine, not a faithful explanation of the output.

Keep reading

related
What Is a Transformer (Neural Network Architecture)?
related
What Is RLHF (Reinforcement Learning from Human Feedback)?
related
What Is an AI Hallucination?
related
What Is Model Evaluation (Benchmarks and Evals)?
referenced by
What Is the Attention Mechanism?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is Fine-Tuning?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.