What Is the Attention Mechanism?
Attention is the operation that lets a model, when computing the representation of one position, draw selectively on every other position — assigning each a learned weight for how relevant it is to the current one. Each token emits a query, and the match between that query and every other token’s key decides how much of each token’s value flows in. It replaced the fixed-size bottleneck of earlier encoders with a direct, weighted view of the whole input.
Queries, keys and values
Each position produces three vectors: a query (what it is looking for), a key (what it offers), and a value (what it contributes if attended to). The weight from one position to another is the similarity of the first’s query to the second’s key, normalised across all positions; the output is the weighted sum of values. Self-attention is the case where queries, keys and values all come from the same sequence — the operation the transformer stacks.
The intuition "attention is interpretable — it shows what the model looks at" is only partly true and is contested in the literature: attention weights correlate with influence but are not a faithful explanation of a prediction. Treating them as a definitive account of model reasoning is a known overreach.
Related standards
Questions
Is attention unique to transformers?
No — it predates them (in recurrent translation models); the transformer’s contribution was to build the whole architecture from attention alone.
Do attention weights explain a model’s decision?
Not reliably. They indicate what was weighted, but research shows they are not a faithful explanation of the output.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.