What Is Tokenization in AI (Subword Tokens)?

In AI, tokenization is splitting text into the units a model actually processes — usually subwords, not whole words or characters. A method like byte-pair encoding learns a vocabulary of frequent fragments, so common words become one token and rare ones split into several. It is distinct from tokenization in crypto, which means representing an asset on a ledger; here a token is a chunk of text, and a model’s context length and cost are measured in these tokens.

Unit
Subword fragments, learned from frequency
Common method
Byte-pair encoding (BPE)
Not to be confused with
Asset tokenization in crypto

Why subwords, and why it leaks into behaviour

Whole-word vocabularies cannot cover every word and handle rare terms badly; character-level units make sequences impractically long. Subword tokenization is the compromise: a fixed vocabulary of fragments that composes any string while keeping frequent words compact. BPE builds that vocabulary by repeatedly merging the most common adjacent pair.

The abstraction is leaky in ways a practitioner must know. A model counts and reasons over tokens, not characters, which is why it miscounts letters in a word and why arithmetic on long numbers is brittle — the digits were split unevenly. Non-English text and code often tokenize less efficiently, costing more tokens for the same content. These are tokenizer artefacts, not reasoning failures.

Related standards

Sennrich et al., 2016 — Neural MT of Rare Words with Subword Units (BPE)

Questions

Why does a model miscount letters in a word?

It sees tokens, not characters — a word may be one token, so the individual letters are not separately represented.

Is this the same as crypto tokenization?

No — unrelated. Here a token is a unit of text; in crypto it is a representation of an asset on a ledger.

Keep reading

related
What Is a Transformer (Neural Network Architecture)?
related
What Is an Embedding?
related
Managing an Agent’s Context Window
related
What Is Semantic Search?
referenced by
Transformer vs RNN: What Changed?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is Fine-Tuning?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.