What Is a Transformer (Neural Network Architecture)?
The transformer is a neural-network architecture that processes a whole sequence at once through stacked layers of self-attention, rather than stepping through it token by token as earlier recurrent networks did. Introduced in 2017, its decisive property is parallelism: because every position attends to every other simultaneously, training scales across modern hardware in ways recurrence could not, which is what made today’s large language models economically trainable.
Why parallelism, not accuracy, was the breakthrough
Recurrent networks read a sequence in order, so each step depends on the last and training cannot be parallelised across the sequence; long-range dependencies also decay. The transformer replaces recurrence with attention, where every token’s representation is computed from a weighted view of all tokens at once. The accuracy gains matter, but the architectural point is that the computation is parallel — the reason model and data scale became practical, and the reason scaling laws could be exploited at all.
The known cost is that self-attention is quadratic in sequence length: doubling the context roughly quadruples the attention compute. Much subsequent research — sparse, linear and flash attention — is an attempt to soften that quadratic without losing what made attention work, and it remains an active, unsettled area.
Related standards
Questions
Are all large language models transformers?
Almost all current ones are, though research into alternatives (state-space models like Mamba) is active; none has displaced the transformer at scale yet.
Is a transformer the same as a large language model?
No — the transformer is the architecture; an LLM is a large model, usually transformer-based, trained on text.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.