⚙️ · Operate

llama.cpp

A portable C and C++ inference engine that runs quantized LLMs on CPUs and edge devices.

In one line

llama.cpp runs quantized GGUF models efficiently on ordinary CPUs and edge hardware, so inference can happen anywhere without a datacenter GPU.

ConceptWhat it is

llama.cpp is an open-source inference engine written in portable C and C++ that runs large language models on commodity hardware, from a laptop CPU to an Apple Silicon Mac to a single-board computer, by loading weights in the GGUF file format. It exists because most serving stacks assume a rented datacenter GPU, while a large class of use cases needs a model that runs locally, offline, or on the edge instead.

Its core move is aggressive quantization — storing weights at 4-, 5-, or 8-bit precision so a multi-billion-parameter model fits in ordinary RAM — paired with hand-tuned kernels for CPU vector units and optional GPU offload of some layers. The design bet is that running anywhere beats running fastest at server scale, which makes it the default engine behind many local-model tools.

How it worksThe mechanics

You download or convert a model into a GGUF file at a chosen quantization level, and llama.cpp memory-maps that file, loads the tokenizer and the computation graph, and executes the transformer forward pass using SIMD-optimized CPU kernels; if a compatible accelerator is present you can offload a configurable number of layers to it through Metal on Apple Silicon or CUDA, ROCm, or Vulkan elsewhere, keeping the rest on CPU. Generation is exposed through the command line, a bundled lightweight HTTP server that speaks an OpenAI-compatible API, or language bindings, and a key-value cache holds prior tokens so each new token reuses past attention state.

At a glanceSee it

llama.cpp diagram
llama.cpp diagram 1

How GGUF shrinks weights — blocks share a scale, k-quants add super-block scales to guard the bits that matter, and over-aggressive quantization surfaces as rising perplexity.

llama.cpp diagram 2

The autoregressive decode loop — prefill fills the KV cache once, each step reuses it to emit one token, and re-reading every weight per token is why CPU inference stays memory-bandwidth bound.

When to use itWhere it fits

  • Local or on-device inference where user data must never leave the machine, such as desktop apps or private notes.
  • Edge and offline deployments — kiosks, field devices, air-gapped environments — with no reliable network or GPU.
  • Prototyping and personal use on a laptop, where spinning up a GPU server is overkill for the volume needed.
  • Cost-sensitive, low-to-moderate traffic where a per-token API bill or an idle GPU is hard to justify.

When NOT to use itLimits & anti-patterns

  • High-concurrency production serving, where continuous batching on GPUs delivers far more tokens per dollar and second.
  • Latency-critical workloads at scale that need the throughput of PagedAttention-style GPU engines like vLLM or TGI.
  • Serving the largest frontier-scale models at full precision, which exceed what CPU or single-device memory can hold.
  • Teams that want a fully managed endpoint and would rather not own quantization, packaging, or client hardware variance.

Trade-offsAdvantages & costs

Advantages
  • Runs practically anywhere — CPU, laptop, phone-class chips, and consumer GPUs — with minimal dependencies.
  • Built-in quantization shrinks models enough to fit capable 7-to-8-billion-parameter models in ordinary RAM.
  • No per-token API cost and no network dependency once the weights are on the device.
  • Broad ecosystem: it is the engine underneath many popular local-model tools, so GGUF weights are widely available.
Trade-offs & costs
  • Lower throughput at server scale than GPU-native engines built around large-batch continuous scheduling.
  • Quantization trades some output quality for size, more noticeable on reasoning-heavy tasks at aggressive bit depths.
  • Per-device performance varies widely with CPU, memory bandwidth, and accelerator, complicating capacity planning.
  • Managing GGUF conversion, quant levels, and distribution across client hardware adds packaging and support work.

ExampleIn the real world

A team building a desktop note-taking tool wants on-device summarization so a customer's text never leaves their machine. They ship an 8-billion-parameter open model quantized to 4-bit GGUF and embed llama.cpp directly in the app. On a mainstream laptop it generates a handful of tokens per second on CPU alone, and on an Apple Silicon Mac it offloads layers to Metal for a clear speedup — with no API bill, no network round-trip, and no GPU server to operate. When the same company later needs a high-traffic cloud summarization endpoint, they keep llama.cpp on the desktop but reach for a batched GPU serving stack in the datacenter, using each engine where its economics fit.

ToolsHow to implement it

  • GGUFthe quantized weight format llama.cpp loads, with widely shared prebuilt weights on the Hugging Face Hub.
  • Ollamaa popular local-model runner built on llama.cpp that packages models and exposes a simple local API.
  • LM Studioa desktop app for downloading and chatting with GGUF models, using llama.cpp under the hood.
  • llama-cpp-pythonPython bindings for embedding the engine in scripts, RAG pipelines, and backends.

Cost & effortWhat it takes

The engine itself is free and open source, and the dominant cost is client or edge hardware you likely already own rather than rented GPUs, so per-inference marginal cost trends toward zero once weights are local. The effort profile is front-loaded: choosing a model and quantization level, validating quality after quantizing, and testing across the range of devices you must support. It shines when volume is modest or must run offline; as steady traffic climbs, the throughput ceiling means a GPU-batched stack usually becomes the cheaper path per token.

A living map of modern AI — kept current every morning