Quantization trades a small amount of accuracy for large gains in speed, memory, and cost.
ConceptWhat it is
Quantization is the process of reducing the numerical precision of a model's weights and activations, typically from 16-bit floating point down to 8-bit or 4-bit integers, to shrink memory footprint and speed up inference. It exists because full-precision models are expensive to serve, and most of that precision is unnecessary for acceptable output quality.
The trade-off is a small, usually measurable, drop in accuracy in exchange for large gains in throughput, memory use, and cost.
How it worksThe mechanics
A quantization algorithm like GPTQ, AWQ, or bitsandbytes maps each weight's continuous range of values to a small set of discrete levels, storing scale factors to reconstruct approximate original values during computation, which lets a GPU hold and compute more of the model in less memory with fewer bits moved per operation.
At a glanceSee it
Zoom into the map-to-fewer-levels step — a scale factor and zero point convert each float into an int bucket and back, while a handful of outlier weights quietly steal precision from all the rest.
Which quantization method to reach for is a real branching decision — whether you can retrain and whether you hold calibration data pick between quantization-aware training, post-training, and dynamic.
When to use itWhere it fits
- Deploying large models on limited GPU memory or consumer hardware.
- High-throughput serving where lower latency per token directly reduces cost.
- Edge or on-device inference where memory is tightly constrained.
- Scaling a self-hosted fleet where every GPU saved is real cost saved.
When NOT to use itLimits & anti-patterns
- Tasks requiring maximum precision, like complex multi-step reasoning or math, where accuracy loss compounds.
- Small models already comfortably fitting on available hardware, where quantization adds complexity for little gain.
Trade-offsAdvantages & costs
Advantages
- Cuts memory footprint by roughly 2 to 4 times, letting bigger models fit on smaller GPUs.
- Increases throughput and lowers per-token latency.
- Reduces serving cost proportionally to the memory and compute saved.
Trade-offs & costs
- Introduces some accuracy degradation, more noticeable on reasoning-heavy tasks.
- Not all quantization methods are supported by all serving frameworks or hardware.
- Requires re-validation of quality after quantizing, adding testing overhead.
ExampleIn the real world
Meta ships official 4-bit and 8-bit quantized versions of its Llama models specifically so developers can run capable models on a single consumer GPU instead of requiring a multi-GPU server.
ToolsHow to implement it
- GPTQpopular post-training quantization method for 4-bit weights.
- AWQactivation-aware quantization preserving accuracy on salient weights.
- bitsandbyteswidely used library for 8-bit and 4-bit quantization in Hugging Face pipelines.
- llama.cppefficient quantized inference for CPU and consumer GPU deployment.
Cost & effortWhat it takes
Quantization itself is a one-time offline process taking minutes to hours; the payoff is ongoing, often 40 to 70 percent lower serving cost per token with modest engineering effort to validate quality.
What changedWhat changed here
Updated this page A reported DeepSeek 4.1 Flash cuts memory use 4x versus DeepSeek 4.0 Flash, which would lower self-hosting cost and allow longer contexts on the same hardware.
Update the DeepSeek V4 Flash page to note the reported DeepSeek 4.1 Flash release and its claimed 4x memory reduction, since the page currently presents V4 Flash as the current Flash-class entry.
Updated this page Quantizing the KV cache can flip which experts a MoE model routes tokens to, because routing is discontinuous.
On sub/pipe-mixture-experts-routing, note that MoE routing is discontinuous and that quantizing the KV cache can flip which experts fire.
Three kinds of claim, strongest first. Signal runs every morning.