Quantization shrinks a model's numbers so it runs cheaper with barely any quality loss.
DefinitionWhat it means
Quantization reduces the numerical precision of a trained model's weights, for example from 16-bit floats down to 4-bit or 8-bit integers, shrinking memory footprint and speeding up inference at the cost of a small, usually acceptable, drop in output quality.
Why it mattersWhy you should care
Quantization is often the difference between a model that fits on affordable hardware or a single GPU and one that does not, so it directly shapes deployment cost and which models can even run on-device or at the edge, making it a standard lever product teams pull before scaling inference.
At a glanceSee it
The upstream decision the linear pipeline hides — pay a retraining budget for quantization-aware training and keep accuracy, or take fast post-training quantization and risk an outlier-driven accuracy cliff.
What actually happens to each number — calibration derives a scale and zero point that snap floats onto an integer grid and are reused to dequantize at matmul, leaving a small error that activation outliers can amplify.
Where you see itIn the wild
- Open-weight model releases shipping quantized GGUF or AWQ variants.
- On-device and edge deployments needing a smaller memory footprint.
- Design discussions on quality versus cost tradeoffs of quantization.