⚙️ · Operate

TensorRT-LLM

NVIDIA's compiled-kernel inference library that squeezes maximum LLM serving speed from its own GPUs.

In one line

TensorRT-LLM compiles a model into a hardware-specific engine of fused CUDA kernels to deliver the lowest latency and highest throughput achievable on NVIDIA GPUs.

ConceptWhat it is

TensorRT-LLM is NVIDIA's open-source inference library that turns a large language model into a compiled TensorRT engine — a serialized bundle of fused, hardware-specific CUDA kernels — instead of running the model through a general-purpose framework at request time. It exists because generic PyTorch execution leaves substantial GPU performance unclaimed, and extracting the last increment of speed from expensive accelerators requires ahead-of-time compilation tuned to the exact chip.

The build step applies kernel fusion, quantization such as FP8, INT8, and INT4/AWQ, a paged KV cache, and tensor or pipeline parallelism, then serves the result — usually behind NVIDIA Triton Inference Server — with in-flight batching. The payoff is class-leading latency and throughput on NVIDIA hardware; the cost is that the engine is NVIDIA-only and pinned to the GPU architecture and shape profile it was compiled for.

How it worksThe mechanics

You start from model weights, convert the checkpoint into TensorRT-LLM's format, then run the engine builder, which fuses attention and feed-forward kernels, applies the chosen quantization, selects the fastest kernel tactics for the target GPU, and serializes everything into an engine bound to a specific architecture and a range of batch sizes and sequence lengths. At serve time a runtime such as Triton loads that engine and drives it with in-flight batching and a paged KV cache, so many concurrent requests share each decode step; any change to the GPU model or the maximum shapes means rebuilding the engine from scratch.

At a glanceSee it

TensorRT-LLM diagram
TensorRT-LLM diagram 1

In-flight batching keeps the GPU saturated by admitting new requests into freed slots on every decode step, instead of draining a whole batch before starting the next.

TensorRT-LLM diagram 2

Choosing engine precision is a real trade among accuracy budget, GPU generation, and whether decode is memory-bound or compute-bound — from FP16 down to weight-only INT4.

When to use itWhere it fits

  • Latency-sensitive, high-throughput production serving on an NVIDIA GPU fleet you already run.
  • Squeezing maximum tokens per second per dollar out of expensive H100-class accelerators.
  • Deployments already standardized on NVIDIA Triton, where TensorRT-LLM is the fastest backend.
  • Serving popular open-weight models such as Llama, Mistral, or Qwen at scale, where quantization gains compound.

When NOT to use itLimits & anti-patterns

  • Any non-NVIDIA target — AMD, Google TPU, AWS Inferentia, or Apple silicon — which it cannot run on.
  • Rapid experimentation or frequently swapped models, where the ahead-of-time build loop slows every iteration.
  • Low, spiky, or early-stage traffic that never justifies the build complexity and GPU commitment.
  • Teams without CUDA and MLOps depth to manage engine builds, versioning, and rebuilds on hardware changes.

Trade-offsAdvantages & costs

Advantages
  • Class-leading latency and throughput on NVIDIA GPUs, typically ahead of general-purpose servers.
  • Deep quantization support (FP8, INT8, INT4, AWQ) that cuts memory and boosts speed with modest quality loss.
  • In-flight batching and a paged KV cache keep GPUs highly utilized under concurrent load.
  • Tight integration with Triton and the broader NVIDIA software stack for production serving.
Trade-offs & costs
  • NVIDIA-only, a hard vendor lock-in to a single hardware ecosystem.
  • Ahead-of-time engine builds are complex and slow, and must be redone when the GPU or shapes change.
  • Engines are pinned to one GPU architecture and shape profile, reducing portability and flexibility.
  • Steeper learning curve than a drop-in server, with more moving parts to operate and debug.

ExampleIn the real world

A team running a customer-support assistant on a fleet of NVIDIA H100s serves a fine-tuned Llama 3 model to thousands of concurrent chats. Under vanilla PyTorch the GPUs sit underutilized and tail latency drifts high, so they convert the checkpoint with TensorRT-LLM, build an FP8-quantized engine tuned for the H100 and their expected sequence lengths, and serve it behind Triton with in-flight batching. Throughput per GPU rises markedly and p99 latency tightens, letting them serve the same traffic on fewer GPUs; the trade is that when they later add a different GPU type, each new architecture requires its own engine rebuild.

ToolsHow to implement it

  • TensorRT-LLMthe library and engine builder itself, NVIDIA's compiled-kernel LLM inference stack.
  • NVIDIA Triton Inference Serverthe production serving layer whose tensorrtllm backend runs the built engines.
  • TensorRTthe underlying general-purpose deep-learning inference compiler that TensorRT-LLM extends.
  • vLLM and Hugging Face TGIcommon non-compiled serving alternatives worth benchmarking against.

Cost & effortWhat it takes

The software is free and open source, so the real cost is the NVIDIA GPUs it runs on — often several dollars to tens of dollars per GPU-hour for data-center accelerators — plus meaningful engineering effort to set up conversion, tune the build, and maintain engine rebuilds as models and hardware evolve. It pays off when sustained, high-utilization serving turns the extra throughput per GPU into fewer GPUs and a lower cost per token; for low or bursty volume, that effort rarely earns back its complexity.

A living map of modern AI — kept current every morning