What Is Chain-of-Thought Prompting?
Chain-of-thought is prompting a model to produce intermediate reasoning steps before its final answer, rather than answering directly. On multi-step problems — arithmetic, logic, planning — this materially improves accuracy, and the effect appears mainly in sufficiently large models. The working explanation is that generating the steps lets the model allocate more computation to the problem and condition each step on the last, approximating a worked solution rather than a single-shot guess.
More computation, not necessarily honest introspection
The practical mechanism is that each generated step becomes context for the next, so the model spends more forward passes on a hard problem and builds its answer incrementally. This is why it helps most where a problem genuinely decomposes into steps, and little on single-fact recall.
A research-grade caveat matters here: the written chain is not guaranteed to be the model’s actual reasoning. Studies show a model can produce a plausible rationale that does not reflect what drove its answer, and can even rationalise a biased conclusion. For an agent acting on its own output, a chain-of-thought is a useful artefact to inspect, not a trustworthy audit of why it decided.
Related standards
Questions
Does chain-of-thought work on any model?
The benefit is strongest in larger models; small models often gain little, and the original work framed it as an emergent behaviour of scale.
Is the reasoning trace trustworthy?
Not inherently — it can be post-hoc rationalisation. It is evidence to examine, not a faithful explanation of the output.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.