Home › Deployment, Inference & LLMOps › Key term › Quantization
Key term · Operate

Quantization

Compressing model weights to smaller numbers to cut memory and cost.

In one line

Quantization shrinks a model's numbers so it runs cheaper with barely any quality loss.

DefinitionWhat it means

Quantization reduces the numerical precision of a trained model's weights, for example from 16-bit floats down to 4-bit or 8-bit integers, shrinking memory footprint and speeding up inference at the cost of a small, usually acceptable, drop in output quality.

Why it mattersWhy you should care

Quantization is often the difference between a model that fits on affordable hardware or a single GPU and one that does not, so it directly shapes deployment cost and which models can even run on-device or at the edge, making it a standard lever product teams pull before scaling inference.

At a glanceSee it

Quantization diagram
Quantization diagram 1

The upstream decision the linear pipeline hides — pay a retraining budget for quantization-aware training and keep accuracy, or take fast post-training quantization and risk an outlier-driven accuracy cliff.

Quantization diagram 2

What actually happens to each number — calibration derives a scale and zero point that snap floats onto an integer grid and are reused to dequantize at matmul, leaving a small error that activation outliers can amplify.

Where you see itIn the wild

  • Open-weight model releases shipping quantized GGUF or AWQ variants.
  • On-device and edge deployments needing a smaller memory footprint.
  • Design discussions on quality versus cost tradeoffs of quantization.
A living map of modern AI — kept current every morning