AI This article was created with the help of AI.
The Problem Each Serving Layer Solves for GPU Inference
When deploying large language models on an existing Kubernetes cluster, a raw inference engine like vLLM or TensorRT-LLM is not enough. An engine handles token generation on a single set of GPUs, but it does not manage cluster-level lifecycle, queue-aware replication, or cross-node routing. Placing a control plane in front of your GPU workers determines how your cluster reacts to traffic spikes, how efficiently your VRAM is utilized, and how much operational glue code your platform engineering team must maintain. Comparing KServe, Ray Serve, and llm-d reveals three distinct architectural abstractions rather than drop-in replacements: llm-d's own founding proposal recommends KServe for teams with high numbers of model deployments or many teams needing distinct model deployments, and states that llm-d does not orchestrate inference workloads directly, instead exposing large-model optimizations into KServe's LLMInferenceService custom resource.
In a vanilla Kubernetes setup, teams often start with standard Deployments and horizontal pod autoscalers (HPA) targeting GPU metrics. This pattern fails quickly for LLMs: GPU duty cycle is a misleading scaling signal, container spin-up times on heavy weights create severe cold-start penalties, and generic round-robin load balancers destroy the KV-cache locality that modern inference runtimes depend on to keep time-to-first-token (TTFT) low. Each framework addresses a different slice of this operational bottleneck.
| Framework | Primary Abstraction Layer | Core Responsibility | What It Leaves to the Operator |
|---|---|---|---|
| KServe | Kubernetes Custom Resources (CRDs) | Declarative model lifecycle, rollout governance, and scale-to-zero | KV-cache routing, disaggregated prefill/decode scheduling, and deep engine telemetry |
| Ray Serve | Python-native distributed actor system | Multi-model dynamic pipelines, complex routing, and request queuing | Kubernetes cluster lifecycle (requires KubeRay) and native container governance |
| llm-d | Distributed LLM scheduling and cache fabric | Prefix-aware routing, prefill/decode disaggregation, and distributed KV management | General model lifecycle, protocol exposure, and infrastructure provisioning |
Understanding these architectural boundaries prevents teams from trying to force a general-purpose model server to solve distributed attention caching, or introducing an entire distributed computing runtime when they only need declarative container autoscaling.
KServe: Standardizing Kubernetes-Native Model Autoscaling
KServe operates as a cloud-native model orchestration control plane, exposing declarative Custom Resource Definitions (CRDs) such as InferenceService and the newer LLMInferenceService, the CRD into which llm-d exposes its large-model optimizations. For platform teams entrenched in GitOps workflows, KServe standardizes how models are packaged, rolled out via blue-green or canary revisions, and exposed through standardized protocols.
Autoscaling Mechanics with Knative and KEDA
KServe does not implement its own autoscaler: it delegates scaling decisions to Knative's Pod Autoscaler or to Kubernetes Event-driven Autoscaling (KEDA). Knative provides concurrency- and request-rate-based scaling alongside native scale-to-zero capabilities by parking ingress requests inside an activator proxy while a new GPU worker spins up. With KEDA, an InferenceService can scale on external metrics such as vLLM's number of waiting requests or KV-cache usage, so scale-out is triggered by queue depth or active inference requests rather than raw compute utilization.
- Declarative GitOps alignment: Model deployments, resource limits, and canary splits live entirely within standard Kubernetes manifests.
- Protocol standardization: Enforces uniform v1/v2 Open Inference Protocol or OpenAI-compatible interfaces across heterogeneous backends.
- Scale-to-zero efficiency: Knative queues incoming traffic during cold starts, preventing request drops when scaling from zero idle replicas.
- Cache-blind routing limits: Standard ingress proxies (such as Envoy or Istio) distribute traffic evenly across replicas, causing frequent KV-cache misses across GPU workers.
While KServe excels at operational governance and infrastructure lifecycle, its networking layer remains largely cache-blind. Standard ingress routers treat every request to a multi-replica LLM pool identically, shuffling multi-turn conversations across random pods and eliminating the throughput benefits of local prefix caching.
Ray Serve: Python-Native Orchestration Beyond the Cluster
Ray Serve takes an application-first approach to model serving by building directly on top of the Ray distributed actor framework. Instead of managing individual containers through Kubernetes manifests, engineers define their serving graph in Python. A deployment consists of Ray actors scheduled across CPU and GPU worker pools, with per-replica resources declared through ray_actor_options and support for fractional accelerators so several light models can share one GPU, making it a natural fit for complex pipelines that combine tokenization, embedding generation, vector search, reranking, and language generation within a single execution graph.
The Dual-Autoscaler Dynamic under KubeRay
Autoscaling in Ray Serve operates at two distinct layers. At the application layer, the Serve autoscaler reacts to traffic spikes by monitoring queue sizes and adding or removing replicas to hold the average number of ongoing requests per replica near a configured target (target_ongoing_requests). If the Ray Autoscaler underneath determines there are not enough available CPUs or GPUs to place those replica actors, it responds by requesting more Ray nodes, and the cloud provider then adds them. Running Ray Serve on Kubernetes requires deploying KubeRay, which bridges Ray cluster specifications to underlying Kubernetes pods.
- Python-native control flow: Write complex dynamic routing, fallback logic, and ensemble pipelines directly in Python code without separate orchestration services.
- Granular resource slicing: Ray actors can reserve fractional GPUs or share accelerator memory across multiple light models without container boundaries.
- Two-tier scaling latency: Scaling up involves two hops: the Serve Autoscaler schedules an actor, and if VRAM is exhausted, KubeRay must request a new pod, compounding scheduling latency.
- Operational divergence: Ray manages its own state, networking, and task scheduling, bypassing standard Kubernetes debugging tools and telemetry pipelines.
Ray Serve provides immense flexibility for data science teams that prioritize rapid iteration in Python over strict infrastructure constraints. However, running a distributed runtime inside another distributed orchestrator increases operational surface area and troubleshooting overhead.
llm-d: Adding Disaggregated Serving and LLM-Specific Routing
llm-d is an open-source distributed inference serving stack for Kubernetes. Launched in May 2025 by Red Hat, Google Cloud, and IBM Research with founding contributors NVIDIA and CoreWeave, it was accepted as a CNCF Sandbox project in March 2026 to address the cluster-level scaling bottlenecks specific to large language models. Where KServe handles general container lifecycle and Ray Serve manages Python execution graphs, llm-d focuses strictly on distributed attention, KV-cache locality, and phase-separated workload orchestration on top of engines like vLLM.
Prefix-Aware Scheduling and Phase Disaggregation
Modern inference efficiency depends on reusing computed KV caches for shared system prompts, multi-turn chat history, and document contexts. Standard load balancers break cache locality. llm-d introduces an intelligent Router and Endpoint Picker that indexes KV-cache block distribution across running vLLM workers, routing requests directly to the replica that already holds matching prefix tokens. Furthermore, llm-d enables prefill/decode (P/D) disaggregation, making the placement decision in the scheduler and using NIXL to abstract high-performance point-to-point KV-cache transfer between the prefill and decode workers.
- Prefix-aware routing: llm-d reports 3x higher output throughput and 2x faster TTFT with prefix-cache-aware routing versus round-robin balancing, measured on Llama 3.1 70B across four AMD MI300X GPUs.
- Prefill and decode disaggregation: up to 70% higher tokens/sec versus standard vLLM in llm-d's reported benchmark of GPT-OSS on NVIDIA B200 instances.
- Distributed cache tiering: Offloads cold KV blocks to host CPU memory or local SSDs, extending effective context capacity beyond physical HBM limits.
- Narrow scope: llm-d does not manage raw pod lifecycle or hardware provisioning; it integrates with higher-level operators like KServe via LLMInferenceService to handle container rollouts.
By addressing attention caching and execution phase separation directly at the network routing layer, llm-d transforms multiple independent vLLM replicas into a unified, high-utilization distributed inference cluster.
The Engineering Costs of Maintaining an Inference Control Plane
Selecting a serving layer is as much an operational commitment as an architectural choice. Each framework automates specific parts of autoscaling GPU inference, but the techniques that make large-scale serving efficient remain, in llm-d's own words, sparsely implemented and a source of high operational toil for the platform teams that run them. In production, that toil takes the form of infrastructure glue code your platform team must still design, test, and maintain.
The Unresolved Operational Responsibilities
None of these frameworks provide a turnkey platform. In production, engineering teams must maintain custom monitoring pipelines to bridge GPU metrics, engine telemetry, and horizontal scaling controllers. When high-concurrency spikes hit, standard rate limiters fail to account for token length variance, requiring custom flow-control proxies to prevent memory exhaustion and CUDA OOM crashes across the fleet.
- Telemetry and metrics aggregation: Standard Prometheus scrapers struggle with high-frequency vLLM metrics. Platform teams must build custom collectors to export queue depth and cache hit rates into autoscaling triggers.
- Flow control and backpressure: None of the base frameworks natively protect upstream clients from token exhaustion during sudden bursts without custom admission controllers.
- Cold-start image distribution: Pulling 20GB to 140GB model weights onto newly spawned GPU nodes requires maintaining localized OCI registries, pre-warmed daemonsets, or shared network file systems.
- Runtime patching and CUDA upgrades: Maintaining underlying container images across CUDA minor versions, driver updates, and flash-attention kernel patches requires continuous validation.
The real cost of a self-hosted control plane is rarely the compute itself: it is the ongoing senior engineering time required to debug distributed state, tune autoscaling timing parameters, and maintain high availability under fluctuating traffic.
Choosing the Right Serving Layer for Your Infrastructure
Choosing between KServe, Ray Serve, and llm-d depends on your team's existing infrastructure maturity, your model architecture complexity, and your latency requirements. These tools can also be combined: for instance, using KServe's LLMInferenceService as the Kubernetes-native lifecycle manager with llm-d handling routing and KV-cache orchestration underneath.
| Requirement / Characteristic | KServe | Ray Serve | llm-d |
|---|---|---|---|
| Primary Audience | Kubernetes platform engineers | Data science and ML application engineers | High-scale LLM infrastructure engineers |
| Configuration Format | Declarative Kubernetes YAML / CRDs | Python code and deployment configurations | Kubernetes Gateway API and routing specs |
| Complex Dynamic Pipelines | Requires external workflow engines | Native support for multi-model DAGs | Focused strictly on LLM request routing |
| KV-Cache Awareness | Requires integration with llm-d or custom proxies | Custom routing actors required | Native prefix-cache indexing and routing |
| Prefill/Decode Separation | Delegated to underlying runtime | Custom actor separation logic required | Native scheduler-directed RPC disaggregation |
Choose KServe if your team operates strict Kubernetes-native environments, relies on GitOps pipelines, and serves a diverse catalog of traditional ML and generative models. Choose Ray Serve if your product relies on complex multi-model pipelines where Python-native flexibility outweighs Kubernetes-native ergonomics. Adopt llm-d when you are serving high-throughput LLM workloads at scale and need to maximize token throughput and minimize TTFT through prefix caching and prefill/decode disaggregation.
When to Bypass Orchestration with a Managed Inference Endpoint
Operating a production-grade inference control plane on Kubernetes requires significant capital and engineering overhead. When calculating the true inference cost per token, teams must account for idle GPU capacity, cluster management overhead, and the continuous maintenance of ingress routers, autoscalers, and storage backends.
Evaluating the Self-Hosted Break-Even Point
For many teams, building and maintaining custom KServe, Ray Serve, or llm-d stacks is an unnecessary distraction from core product development. If your workload exhibits variable or bursty traffic patterns, paying for dedicated GPU nodes to keep baseline replicas warm results in high cost waste. In these scenarios, consuming open-source models through a fully managed, consumption-based endpoint eliminates the operational burden entirely.
- Zero control plane overhead: Eliminate the need to configure Knative, debug KubeRay actors, or tune distributed KV caches.
- True consumption pricing: Pay strictly for the input and output tokens you consume without paying for idle GPU provisioning time.
- Full SDK compatibility: Integrate directly with standard OpenAI-compatible libraries without maintaining custom ingress gateways.
- Infrastructure control when needed: For workloads that justify a dedicated control plane, provisioning raw compute on dedicated nodes provides full kernel and driver access.
At Lyceum, we provide an EU-sovereign AI cloud built to give European engineering teams flexible options across the infrastructure stack. For teams seeking to eliminate orchestration complexity, Lyceum Serverless Inference provides high-performance, per-token access to open-source models with zero data retention and no egress fees. For platform teams committed to maintaining their own Kubernetes control plane, Lyceum delivers raw on-demand GPU VMs and dedicated GPU clusters with per-second billing, providing the unconstrained hardware performance required to run your own custom inference stack.