RAG vs Long Context: How Should a Model Get Its Knowledge?

Both give a model knowledge it was not trained on; they differ in how. RAG retrieves only the relevant passages and places those in the context. A long-context approach puts the whole corpus in the window and lets the model attend over all of it. Long context is simpler and avoids retrieval errors, but costs grow with the tokens in the window and relevant facts can get lost in the middle. RAG stays cheap and current at the price of a retrieval that can miss.

RAG
Retrieve the relevant bit, then generate
Long context
Put everything in the window
Trade
Retrieval risk vs token cost and lost-in-the-middle

Not either/or

Longer context windows did not kill RAG, as was briefly predicted. Attention cost grows with window size, so stuffing a large corpus into every call is expensive at scale; and models attend unevenly across a long window, with documented degradation for facts placed in the middle. RAG keeps each call small and the knowledge updatable without retraining, but inherits whatever its retriever misses.

The mature pattern uses both: retrieve to narrow a large corpus to a strong candidate set, then rely on a capable context window to reason over that set. The useful question is not which wins but how much to retrieve versus how much to place directly — an engineering trade, not a doctrine.

Related standards

Lewis et al., 2020 — Retrieval-Augmented Generation

Questions

Did long context windows make RAG obsolete?

No. Token cost scales with the window and models lose facts in the middle of long inputs; RAG remains cheaper and more current for large corpora.

Can you use both together?

Yes, and serious systems do — retrieve to narrow the corpus, then reason over the result in context.

Keep reading

related
Retrieval-Augmented Generation (RAG) for Agents
related
What Is Semantic Search?
related
Managing an Agent’s Context Window
related
What Is a Vector Database?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is Fine-Tuning?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is a Transformer (Neural Network Architecture)?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.