RAG vs Fine-Tuning: Which Should You Use?

They solve different problems and are often used together. Retrieval-augmented generation leaves the model’s weights fixed and supplies knowledge at query time by fetching relevant passages — best for facts that change and claims that must cite a source. Fine-tuning changes the weights to teach behaviour, format, or a narrow skill — best for how the model should act, not what it should currently know. Use retrieval for knowledge, fine-tuning for behaviour, and both when you need each.

RAG
Knowledge at query time, weights unchanged
Fine-tuning
Behaviour baked into the weights
Rule of thumb
Facts → RAG; behaviour → fine-tune

Knowledge versus behaviour

The cleanest way to choose is to ask whether the thing you want to add is a fact or a behaviour. A changing fact — today’s price, this customer’s history, a policy that was updated last week — belongs in retrieval, where it can be fresh and cited. A durable behaviour — always return this JSON shape, adopt this voice, perform this narrow task — belongs in fine-tuning.

Most production systems use both: a capable base model, fine-tuned for how it should behave, retrieving the facts it should reason over. The question is rarely "which" and usually "what goes where".

Related standards

Retrieval-Augmented Generation — Lewis et al., 2020

Questions

Is RAG cheaper than fine-tuning?

Usually to start — it avoids a training run and keeps knowledge updatable — but it adds retrieval infrastructure and per-query cost.

Can fine-tuning replace retrieval for facts?

Poorly. Fine-tuned facts go stale, cannot cite a source, and are costly to update — the reasons retrieval exists.

Keep reading

related
Retrieval-Augmented Generation (RAG) for Agents
related
What Is Fine-Tuning?
related
What Is an Embedding?
related
What Is a Vector Database?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
What Is a Transformer (Neural Network Architecture)?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.