What Is a Transformer (Neural Network Architecture)?

The transformer is a neural-network architecture that processes a whole sequence at once through stacked layers of self-attention, rather than stepping through it token by token as earlier recurrent networks did. Introduced in 2017, its decisive property is parallelism: because every position attends to every other simultaneously, training scales across modern hardware in ways recurrence could not, which is what made today’s large language models economically trainable.

Core operation
Self-attention over the whole sequence
Introduced
Vaswani et al., 2017
Decisive edge
Parallel training, not sequential

Why parallelism, not accuracy, was the breakthrough

Recurrent networks read a sequence in order, so each step depends on the last and training cannot be parallelised across the sequence; long-range dependencies also decay. The transformer replaces recurrence with attention, where every token’s representation is computed from a weighted view of all tokens at once. The accuracy gains matter, but the architectural point is that the computation is parallel — the reason model and data scale became practical, and the reason scaling laws could be exploited at all.

The known cost is that self-attention is quadratic in sequence length: doubling the context roughly quadruples the attention compute. Much subsequent research — sparse, linear and flash attention — is an attempt to soften that quadratic without losing what made attention work, and it remains an active, unsettled area.

Related standards

Vaswani et al., 2017 — Attention Is All You Need

Questions

Are all large language models transformers?

Almost all current ones are, though research into alternatives (state-space models like Mamba) is active; none has displaced the transformer at scale yet.

Is a transformer the same as a large language model?

No — the transformer is the architecture; an LLM is a large model, usually transformer-based, trained on text.

Keep reading

related
What Is the Attention Mechanism?
related
What Is a Mixture-of-Experts Model?
related
Transformer vs RNN: What Changed?
related
What Is Tokenization in AI (Subword Tokens)?
referenced by
What Is Chain-of-Thought Prompting?
referenced by
What Is Model Distillation?
referenced by
What Is Model Quantization?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is Fine-Tuning?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.