RAG vs Fine-Tuning: Which Should You Use?
They solve different problems and are often used together. Retrieval-augmented generation leaves the model’s weights fixed and supplies knowledge at query time by fetching relevant passages — best for facts that change and claims that must cite a source. Fine-tuning changes the weights to teach behaviour, format, or a narrow skill — best for how the model should act, not what it should currently know. Use retrieval for knowledge, fine-tuning for behaviour, and both when you need each.
Knowledge versus behaviour
The cleanest way to choose is to ask whether the thing you want to add is a fact or a behaviour. A changing fact — today’s price, this customer’s history, a policy that was updated last week — belongs in retrieval, where it can be fresh and cited. A durable behaviour — always return this JSON shape, adopt this voice, perform this narrow task — belongs in fine-tuning.
Most production systems use both: a capable base model, fine-tuned for how it should behave, retrieving the facts it should reason over. The question is rarely "which" and usually "what goes where".
Related standards
Questions
Is RAG cheaper than fine-tuning?
Usually to start — it avoids a training run and keeps knowledge updatable — but it adds retrieval infrastructure and per-query cost.
Can fine-tuning replace retrieval for facts?
Poorly. Fine-tuned facts go stale, cannot cite a source, and are costly to update — the reasons retrieval exists.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.