Managing an Agent’s Context Window

A context window is the finite span of tokens a model can attend to at once. Managing it is the art of keeping the right facts in view: summarising completed work, retrieving only what the current step needs, and spilling the rest to a durable store the agent can query again. An agent that treats the window as infinite forgets its own instructions; one that over-trims loses the thread it needs to finish.

Window
The tokens a model attends to at once
Overflow
Summarise, retrieve, spill to a store
Too little
The agent forgets its instructions
Too much
Cost and noise crowd out the signal

The durable-memory escape hatch

The window is working memory, not storage. The reliable pattern is to keep a compact working set in context and push everything else to retrievable memory — a store, a knowledge graph — so the agent can pull a fact back when the step needs it rather than carrying all facts all the time.

Questions

Do bigger windows make this unnecessary?

No. Larger windows raise the ceiling but cost more and dilute attention; retrieval and summarisation stay the economical path for long-running agents.

Where does the spilled context go?

To durable agent memory — a vector store or knowledge graph the agent queries on demand, keeping provenance so a recalled fact can be trusted.

Where this lives in the estate

FlashyOS — context as a first-class, auditable input

Keep reading

related
What Is Agent Context (and the Context Window)?
related
What Is Agent Memory?
related
Retrieval-Augmented Generation (RAG) for Agents
related
What Is an Agent Knowledge Graph?

By Michael Gord · published 2026-10-04 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.