What Is Model Quantization?
Quantization lowers the numerical precision of a model’s weights — from 16-bit floats to 8-bit or 4-bit integers, say — so the model takes less memory and runs faster, with little loss in quality when done carefully. It changes how the existing model is stored and computed, not what it learned. It is one of the main reasons capable models can run on a single GPU or a laptop, which in turn is what makes a large fleet of agents affordable.
Precision you can spend
A weight stored as a 16-bit number carries more precision than inference usually needs. Quantization maps those weights to a coarser grid — 8 or even 4 bits — cutting memory roughly in proportion and speeding up the arithmetic. The art is in doing it without accuracy collapse: outlier weights carry disproportionate information, and the influential LLM.int8() work showed that handling those outliers separately is what makes aggressive quantization viable at large scale.
The trade-off is real and monotonic at the extremes: push precision low enough and quality degrades, unevenly across tasks. Quantization is a cost lever, not a free lunch — the right setting depends on how much quality a use case can spend for the savings.
Related standards
Questions
Does quantization make a model worse?
Mild quantization is often nearly lossless; aggressive quantization degrades quality, and the amount varies by task and method.
Is a quantized model retrained?
Not necessarily — post-training quantization converts an existing model directly; quantization-aware training is a separate, more involved option.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.