Transformer vs RNN: What Changed?

A recurrent neural network processes a sequence one step at a time, carrying a hidden state forward; a transformer processes the whole sequence in parallel through attention. The consequences follow directly: RNNs are sequential and hard to parallelise in training, and long-range signals fade through the chain; transformers train in parallel and connect any two positions directly, at the cost of compute that grows quadratically with length. The shift unlocked training at the scale that produced modern language models.

RNN
Sequential state, fades over distance
Transformer
Parallel attention, any-to-any
Cost moved
From sequential depth to quadratic length

The trade that was worth it

An RNN’s sequential dependence is elegant and memory-light, but it caps how much training can be parallelised and lets distant dependencies decay through many steps — the vanishing-gradient problem that gated units only partly fixed. The transformer pays a different price: attention is quadratic in sequence length, so long contexts are expensive. At the scales that mattered, trading sequential depth for parallelisable quadratic width was decisively the right bargain.

The story is not fully closed. Modern state-space models revisit the recurrent idea with parallelisable training and linear scaling in length, and are a genuine research direction — but as of now none has displaced the transformer at frontier scale.

Related standards

Vaswani et al., 2017 — Attention Is All You Need

Questions

Are RNNs obsolete?

Largely superseded for large-scale language work, though they remain useful in constrained settings; new state-space models revisit their strengths.

Why not just use RNNs for long sequences?

Their signal fades over distance and they cannot be parallelised in training the way transformers can — the reasons they were displaced.

Keep reading

related
What Is a Transformer (Neural Network Architecture)?
related
What Is the Attention Mechanism?
related
What Is a Mixture-of-Experts Model?
related
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is Fine-Tuning?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.