What Is Tokenization in AI (Subword Tokens)?
In AI, tokenization is splitting text into the units a model actually processes — usually subwords, not whole words or characters. A method like byte-pair encoding learns a vocabulary of frequent fragments, so common words become one token and rare ones split into several. It is distinct from tokenization in crypto, which means representing an asset on a ledger; here a token is a chunk of text, and a model’s context length and cost are measured in these tokens.
Why subwords, and why it leaks into behaviour
Whole-word vocabularies cannot cover every word and handle rare terms badly; character-level units make sequences impractically long. Subword tokenization is the compromise: a fixed vocabulary of fragments that composes any string while keeping frequent words compact. BPE builds that vocabulary by repeatedly merging the most common adjacent pair.
The abstraction is leaky in ways a practitioner must know. A model counts and reasons over tokens, not characters, which is why it miscounts letters in a word and why arithmetic on long numbers is brittle — the digits were split unevenly. Non-English text and code often tokenize less efficiently, costing more tokens for the same content. These are tokenizer artefacts, not reasoning failures.
Related standards
Questions
Why does a model miscount letters in a word?
It sees tokens, not characters — a word may be one token, so the individual letters are not separately represented.
Is this the same as crypto tokenization?
No — unrelated. Here a token is a unit of text; in crypto it is a representation of an asset on a ledger.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.