⚙️ · Operate

Ray Serve

A Python-native serving layer that scales and stitches multi-model inference pipelines into one deployable graph.

In one line

Ray Serve turns ordinary Python model code into independently scalable deployments you can compose into a single multi-step inference graph.

ConceptWhat it is

Ray Serve is a framework-agnostic model-serving library built on Ray, the distributed compute engine. It exists to solve two problems that plain containers handle badly: scaling arbitrary Python inference code across a cluster, and composing several models and steps into one pipeline. The unit of work is a deployment — a pool of replica actors wrapping a Python class or function — and deployments wire together into a deployment graph.

Real applications are rarely a single model call: there is preprocessing, retrieval, an embedding model, an LLM, a reranker, and business logic in between. Ray Serve lets each of those be its own deployment with its own replicas, hardware, and autoscaling, connected in plain Python rather than a mesh of microservices. It is deliberately general, so on its own it is not LLM-specialised — for competitive decode throughput you point a deployment at a purpose-built engine like vLLM.

How it worksThe mechanics

You decorate a Python class or function with @serve.deployment, and each becomes a group of replica actors placed on the Ray cluster. Deployments are bound together — one calls another through a deployment handle — to form a graph, then launched with serve run or a declarative Serve config file. Ray Serve schedules replicas across nodes, load-balances requests to them, packs fractional GPUs, batches requests dynamically, and autoscales each deployment independently based on its own in-flight request load, all behind a FastAPI-integrated HTTP or gRPC ingress.

At a glanceSee it

Ray Serve diagram
Ray Serve diagram 1

Inside one replica, dynamic batching accumulates concurrent requests until the batch fills or a timer fires, then serves them all in one vectorized GPU pass — but too long a wait inflates tail latency.

Ray Serve diagram 2

How the Ray scheduler bin-packs fractional-GPU replicas onto shared physical GPUs and requests a new node only when no slot fits — with the caveat that fractional splits share VRAM without hard isolation.

When to use itWhere it fits

  • Multi-step pipelines where each stage — preprocessing, an embedding model, an LLM, a reranker — needs its own scaling and hardware.
  • Python-native teams that want serving logic to stay in Python instead of stitched-together services and YAML.
  • Workloads mixing CPU and GPU stages, where fractional-GPU packing and per-stage autoscaling keep cost proportional to the real bottleneck.
  • When you already run Ray for training or batch and want serving inside the same cluster and ecosystem.

When NOT to use itLimits & anti-patterns

  • A single LLM behind an OpenAI-compatible endpoint, where a dedicated server like vLLM or TGI is simpler and faster out of the box.
  • Small, low-traffic apps where a managed API or one container is far less operational overhead.
  • Teams without the appetite to run and debug a Ray cluster — actors, handles, the object store, placement groups.
  • Latency-critical single-model serving where the extra composition and routing hop earns nothing.

Trade-offsAdvantages & costs

Advantages
  • Composes arbitrary Python models and business logic into one deployable graph, with no glue microservices to build.
  • Each deployment scales and is resourced independently — separate replica counts and fractional GPUs per stage.
  • Framework-agnostic: serve PyTorch, TensorFlow, XGBoost, or a vLLM engine side by side in the same graph.
  • Built on Ray, so serving shares a cluster and tooling with Ray Train, Data, and Tune.
Trade-offs & costs
  • Not LLM-specialised on its own; you must wire in vLLM or Ray Serve LLM to reach competitive decode throughput.
  • Running and debugging a Ray cluster is real operational surface compared with a single serving container.
  • Another distributed system to learn — actors, deployment handles, the object store, and placement.
  • Overkill for a lone single-model endpoint, where a dedicated inference server does more with less setup.

ExampleIn the real world

A team building a document question-answering service models it as one Ray Serve graph: an ingress deployment receives the question, a CPU-only preprocess deployment cleans and chunks it, an embedding deployment on a fraction of a GPU vectorises the query, a retrieval deployment hits the vector store, and a generation deployment running a vLLM engine on full GPUs produces the answer with a reranker in between. When traffic spikes, only the embedding and generation deployments autoscale up more GPU replicas while the CPU stages stay flat, so spend tracks the actual bottleneck rather than the whole pipeline.

ToolsHow to implement it

  • Ray Servethe serving library itself, shipped as part of Ray.
  • Ray Serve LLMLLM-specific deployments with a vLLM backend and an OpenAI-compatible API.
  • KubeRaythe Kubernetes operator for running Ray and Ray Serve clusters in production.
  • vLLMhigh-throughput LLM engine commonly used as the model backend inside a deployment.

Cost & effortWhat it takes

Ray Serve is open source and free; the cost is the cluster underneath and the engineering to run it. You pay for the GPU and CPU nodes, and fractional-GPU packing plus per-deployment autoscaling are the main levers that keep that bill proportional to load. Effort is moderate to high: expressing a pipeline as deployments is quick in Python, but operating a Ray cluster — KubeRay, autoscaler tuning, observability — is a standing commitment that pays off mainly once you have genuine multi-model composition or scale.

A living map of modern AI — kept current every morning