LMDeploy serves open LLMs with built-in low-bit quantization and the TurboMind engine, trading a smaller ecosystem for strong throughput on self-hosted GPUs.
ConceptWhat it is
LMDeploy is a serving and compression toolkit for large language models, built by the InternLM team in the OpenMMLab ecosystem at Shanghai AI Laboratory. It bundles two things most stacks keep separate: a high-performance inference runtime called TurboMind (a hand-optimized C++/CUDA engine descended from FasterTransformer) and a first-class quantization pipeline, so the same tool that shrinks a model also serves it. It exists because self-hosters usually hit a memory wall before a compute wall — a model that will not fit in VRAM cannot be served at any speed — and LMDeploy treats low-bit weights as the default path rather than an afterthought.
Under the hood it uses continuous (persistent) batching and a paged KV cache to keep the GPU busy across many concurrent requests, exposes an OpenAI-compatible HTTP endpoint so existing clients work unchanged, and supports AWQ (W4A16), W8A8, and KV-cache quantization. A second PyTorch engine covers architectures TurboMind has not yet hand-tuned.
How it worksThe mechanics
You point LMDeploy at a Hugging Face checkpoint and, optionally, first run its quantizer to produce 4-bit AWQ weights. TurboMind converts the weights into its own optimized format and loads them onto the GPU, sharding across cards with tensor parallelism when the model is large. Incoming requests join a persistently running batch, with each sequence's tokens allocated into a paged KV cache so prompts of different lengths share GPU memory efficiently instead of reserving worst-case blocks. The engine streams generated tokens back through the OpenAI-compatible server, and as soon as a request finishes its slot is reused by a waiting one, keeping utilization high under concurrency.
At a glanceSee it
The engine fork the serving pipeline hides — LMDeploy routes each model to TurboMind for raw speed or the PyTorch backend for breadth, then serves both behind one API.
Inside the single AWQ box — activation statistics single out the weights worth protecting so the rest can be crushed to 4 bits with little quality loss.
When to use itWhere it fits
- You are self-hosting an open model that will not fit in VRAM at full precision and need 4-bit AWQ or W8A8 to make it deployable.
- You want serving and quantization from one tool rather than gluing a separate quantizer to a separate runtime.
- You are running InternLM, Qwen, Llama, or similar well-supported architectures where TurboMind is already tuned.
- Throughput per GPU dollar matters more to you than breadth of exotic model support.
When NOT to use itLimits & anti-patterns
- You need day-one support for a brand-new or unusual architecture, where a smaller community lags the largest engines.
- You want the deepest ecosystem of integrations, plugins, and community troubleshooting — vLLM or TGI are safer defaults there.
- Accuracy is non-negotiable and even mild quantization loss is unacceptable, in which case serve full precision.
- You are on non-NVIDIA hardware or need a pure-Python stack, since TurboMind's edge is CUDA-specific.
Trade-offsAdvantages & costs
Advantages
- Quantization is built in and treated as a first-class path, so fitting large models onto modest GPUs is straightforward.
- TurboMind delivers strong throughput and low latency on its supported models.
- The OpenAI-compatible API drops into existing client code with no rewrite.
- Continuous batching plus a paged KV cache keeps GPU utilization high under concurrent load.
Trade-offs & costs
- Smaller community and ecosystem than vLLM or TGI, meaning fewer integrations and slower coverage of new models.
- The C++/CUDA TurboMind core is harder to hack or debug than a Python engine when something breaks.
- Aggressive low-bit quantization can degrade output quality and needs evaluation before production.
- The best gains are NVIDIA-specific, so portability across accelerators is limited.
ExampleIn the real world
A team wants to self-host a 72B open model but has GPUs with only 48GB of memory, far short of the roughly 140GB an FP16 copy would need. They run LMDeploy's AWQ quantizer to produce 4-bit W4A16 weights — about a quarter of the footprint — load them into TurboMind, and enable a paged KV cache. The model now fits on a single card with headroom for concurrent sessions, and because the endpoint is OpenAI-compatible their existing chat client connects by changing only the base URL. Under load, continuous batching lets dozens of users share the GPU while first-token latency stays acceptable, turning a deployment that was impossible at full precision into one that runs on hardware they already own.
ToolsHow to implement it
- LMDeploy and its TurboMind engine — the serving core plus AWQ, W8A8, and KV-cache quantization.
- vLLM, Text Generation Inference (TGI), and SGLang — the main alternative open-source serving engines.
- TensorRT-LLM — NVIDIA's lower-level high-performance engine, a common point of comparison.
- AWQ and GPTQ — the low-bit quantization methods that LMDeploy and its peers rely on.
Cost & effortWhat it takes
LMDeploy itself is open source and free; the real cost is the self-hosting operation. You supply and run the GPUs, and the payoff of built-in quantization is directly fewer or smaller cards for a given model. Effort is moderate — standing up a quantized endpoint is a short path, but you own capacity planning, an accuracy-versus-throughput evaluation to validate the quantized weights, and the smaller-community tax of fewer worked examples when you hit an edge case. Compared with a hosted API you trade per-token billing for fixed GPU spend that pays off only at sustained, high-utilization traffic.