RAG vs Long Context: How Should a Model Get Its Knowledge?
Both give a model knowledge it was not trained on; they differ in how. RAG retrieves only the relevant passages and places those in the context. A long-context approach puts the whole corpus in the window and lets the model attend over all of it. Long context is simpler and avoids retrieval errors, but costs grow with the tokens in the window and relevant facts can get lost in the middle. RAG stays cheap and current at the price of a retrieval that can miss.
Not either/or
Longer context windows did not kill RAG, as was briefly predicted. Attention cost grows with window size, so stuffing a large corpus into every call is expensive at scale; and models attend unevenly across a long window, with documented degradation for facts placed in the middle. RAG keeps each call small and the knowledge updatable without retraining, but inherits whatever its retriever misses.
The mature pattern uses both: retrieve to narrow a large corpus to a strong candidate set, then rely on a capable context window to reason over that set. The useful question is not which wins but how much to retrieve versus how much to place directly — an engineering trade, not a doctrine.
Related standards
Questions
Did long context windows make RAG obsolete?
No. Token cost scales with the window and models lose facts in the middle of long inputs; RAG remains cheaper and more current for large corpora.
Can you use both together?
Yes, and serious systems do — retrieve to narrow the corpus, then reason over the result in context.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.