What Is Model Quantization?

Quantization lowers the numerical precision of a model’s weights — from 16-bit floats to 8-bit or 4-bit integers, say — so the model takes less memory and runs faster, with little loss in quality when done carefully. It changes how the existing model is stored and computed, not what it learned. It is one of the main reasons capable models can run on a single GPU or a laptop, which in turn is what makes a large fleet of agents affordable.

Reduces
Numerical precision of the weights
Buys
Less memory, faster inference
Preserves
What the model learned (to a point)

Precision you can spend

A weight stored as a 16-bit number carries more precision than inference usually needs. Quantization maps those weights to a coarser grid — 8 or even 4 bits — cutting memory roughly in proportion and speeding up the arithmetic. The art is in doing it without accuracy collapse: outlier weights carry disproportionate information, and the influential LLM.int8() work showed that handling those outliers separately is what makes aggressive quantization viable at large scale.

The trade-off is real and monotonic at the extremes: push precision low enough and quality degrades, unevenly across tasks. Quantization is a cost lever, not a free lunch — the right setting depends on how much quality a use case can spend for the savings.

Related standards

Dettmers et al., 2022 — LLM.int8()

Questions

Does quantization make a model worse?

Mild quantization is often nearly lossless; aggressive quantization degrades quality, and the amount varies by task and method.

Is a quantized model retrained?

Not necessarily — post-training quantization converts an existing model directly; quantization-aware training is a separate, more involved option.

Keep reading

related
What Is Model Distillation?
related
What Is a Mixture-of-Experts Model?
related
What Is a Transformer (Neural Network Architecture)?
related
What Is Fine-Tuning?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.