Transformer vs RNN: What Changed?
A recurrent neural network processes a sequence one step at a time, carrying a hidden state forward; a transformer processes the whole sequence in parallel through attention. The consequences follow directly: RNNs are sequential and hard to parallelise in training, and long-range signals fade through the chain; transformers train in parallel and connect any two positions directly, at the cost of compute that grows quadratically with length. The shift unlocked training at the scale that produced modern language models.
The trade that was worth it
An RNN’s sequential dependence is elegant and memory-light, but it caps how much training can be parallelised and lets distant dependencies decay through many steps — the vanishing-gradient problem that gated units only partly fixed. The transformer pays a different price: attention is quadratic in sequence length, so long contexts are expensive. At the scales that mattered, trading sequential depth for parallelisable quadratic width was decisively the right bargain.
The story is not fully closed. Modern state-space models revisit the recurrent idea with parallelisable training and linear scaling in length, and are a genuine research direction — but as of now none has displaced the transformer at frontier scale.
Related standards
Questions
Are RNNs obsolete?
Largely superseded for large-scale language work, though they remain useful in constrained settings; new state-space models revisit their strengths.
Why not just use RNNs for long sequences?
Their signal fades over distance and they cannot be parallelised in training the way transformers can — the reasons they were displaced.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.