Home › LLMs & Foundation Models › SLM (small)
🤖 · Models

SLM (small)

Compact language models built to run cheaply and locally without sacrificing task-specific quality.

In one line

Small language models trade some general capability for dramatically lower cost, latency, and on-device deployability.

ConceptWhat it is

A Small Language Model (SLM) is a language model, typically under 10 billion parameters, deliberately optimized for efficiency rather than maximal general capability, often through careful data curation, distillation from a larger teacher model, or architecture tuning. It exists because many real-world tasks do not need a frontier-scale model, and running a smaller model brings major savings in cost, latency, and the ability to run on-device or at the edge.

Well-trained SLMs can match or exceed much larger older models on targeted tasks by using higher-quality training data and techniques like distillation instead of relying purely on scale.

How it worksThe mechanics

SLMs are usually produced either by pretraining a smaller architecture on a carefully curated, high-quality data mix, or by distilling knowledge from a larger teacher model, training the small model to mimic the teacher's outputs or internal representations; techniques like quantization and pruning further shrink the model for deployment on mobile or edge hardware.

At a glanceSee it

SLM (small) diagram
SLM (small) diagram 1

A routing decision — reach for a frontier model only when broad knowledge or hard reasoning is genuinely required, otherwise a small, domain-tuned model wins on cost, latency, and privacy.

SLM (small) diagram 2

A confidence-gated cascade in production — the small model answers most requests cheaply, escalates only the hard ones to a frontier model, and those hard cases feed a distillation loop that keeps sharpening the small model.

When to use itWhere it fits

  • On-device or edge deployment where memory and compute are constrained.
  • High-volume, latency-sensitive tasks where per-request cost must stay low.
  • Narrow, well-defined tasks like classification or extraction that do not need frontier general knowledge.
  • Privacy-sensitive use cases where data cannot leave the device or local network.

When NOT to use itLimits & anti-patterns

  • Open-ended, broad-knowledge tasks requiring deep world knowledge or complex reasoning.
  • Tasks needing strong multilingual or long-context capability beyond the small model's training.
  • Situations where the cost difference versus a frontier API model is not the deciding factor and maximum quality is required.

Trade-offsAdvantages & costs

Advantages
  • Dramatically lower inference cost and latency than frontier-scale models.
  • Can run fully on-device, improving privacy and offline availability.
  • Faster to fine-tune and iterate on for narrow tasks.
  • Lower memory footprint enables deployment on mobile and edge hardware.
Trade-offs & costs
  • Weaker general knowledge and reasoning than frontier-scale models.
  • Can struggle with long context or nuanced multi-step tasks.
  • Quality gap widens on open-ended or unfamiliar tasks outside its training focus.
  • Still requires careful evaluation to confirm it meets the bar for a given task.

ExampleIn the real world

Microsoft's Phi-3 and Google's Gemma models run competitively with much larger models on many benchmarks while being small enough to run on a laptop or phone.

ToolsHow to implement it

  • Microsoft Phi-3small model family tuned via high-quality synthetic data.
  • Google Gemmaopen-weight small model optimized for on-device use.
  • llama.cpp / Ollamarun quantized small models locally with minimal setup.
  • MLC-LLMdeploys small models efficiently on mobile and edge devices.

Cost & effortWhat it takes

Very low inference cost and latency, often runs on CPU or mobile NPU; distillation or fine-tuning is cheap relative to training a large model; needs a well-scoped task to close the quality gap with larger models.

A living map of modern AI — kept current every morning