What Is Prompt Injection?
Prompt injection is an attack that hides instructions in content a model reads — a web page, an email, a tool result — so the model follows the attacker rather than its operator. For an autonomous agent with tools and credentials it is acute: injected text can redirect the agent to exfiltrate data or move funds. There is no complete fix, so defences layer input isolation, least-privilege tools, and a human in the loop for consequential actions.
Why it has no clean fix
A language model does not reliably distinguish instructions it was given from instructions that arrive inside the data it was asked to process — they are all just text. That is a property of how the models work, not a bug to patch, so the content an agent ingests is an attack surface by default.
The estate’s answer is structural: an agent acts under scoped authority with a human consenting to consequential actions, so a successful injection is contained by what the agent was ever permitted to do.
Related standards
Questions
Is prompt injection the same as jailbreaking?
They overlap. Jailbreaking targets a model’s own guardrails; prompt injection targets an application by smuggling instructions through its inputs.
Can it be fully prevented?
Not currently. It is managed by limiting what an agent can do and keeping a human in the loop for anything costly or irreversible.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.