What Is the Attention Mechanism?

Attention is the operation that lets a model, when computing the representation of one position, draw selectively on every other position — assigning each a learned weight for how relevant it is to the current one. Each token emits a query, and the match between that query and every other token’s key decides how much of each token’s value flows in. It replaced the fixed-size bottleneck of earlier encoders with a direct, weighted view of the whole input.

Computes
A weighted combination of all positions
Mechanism
Query–key match sets each weight
First major use
Bahdanau et al., 2015, for translation

Queries, keys and values

Each position produces three vectors: a query (what it is looking for), a key (what it offers), and a value (what it contributes if attended to). The weight from one position to another is the similarity of the first’s query to the second’s key, normalised across all positions; the output is the weighted sum of values. Self-attention is the case where queries, keys and values all come from the same sequence — the operation the transformer stacks.

The intuition "attention is interpretable — it shows what the model looks at" is only partly true and is contested in the literature: attention weights correlate with influence but are not a faithful explanation of a prediction. Treating them as a definitive account of model reasoning is a known overreach.

Related standards

Bahdanau et al., 2015 — Neural Machine Translation by Jointly Learning to Align and Translate

Questions

Is attention unique to transformers?

No — it predates them (in recurrent translation models); the transformer’s contribution was to build the whole architecture from attention alone.

Do attention weights explain a model’s decision?

Not reliably. They indicate what was weighted, but research shows they are not a faithful explanation of the output.

Keep reading

related
What Is a Transformer (Neural Network Architecture)?
related
Managing an Agent’s Context Window
related
What Is Chain-of-Thought Prompting?
related
What Is an Embedding?
referenced by
Transformer vs RNN: What Changed?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is Fine-Tuning?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.