AI This article was created with the help of AI.
The Hardware Imbalance: Prefill vs. Decode
When serving large language models at scale, running prompt ingestion and autoregressive token generation on the same physical GPU creates an immediate structural conflict. These two execution phases, prefill and decode, represent opposite computational workloads with conflicting hardware bottlenecks. Placing both in a shared execution pipeline forces a compromise that degrades system efficiency as context windows grow.
The prefill phase processes the entire input context in parallel. It ingests thousands of prompt tokens simultaneously, creating large matrix multiplications (GEMMs) that fully saturate Tensor Cores. NVIDIA's matrix multiplication background guide works the arithmetic out directly: increasing a GEMM to 8192 x 8192 x 8192 raises arithmetic intensity to 2730 floating-point operations per byte accessed, far above the FLOPS:B ratio of 138.9 the guide assumes for a V100, so the operation is math limited. Prefill is therefore overwhelmingly compute-bound.
In contrast, autoregressive decoding produces exactly one token per sequence per forward step. Each step requires loading every model weight and the accumulated key-value (KV) cache from High Bandwidth Memory (HBM) into SRAM just to compute a single token vector. The same NVIDIA guide concludes from that arithmetic that matrix-vector products (general matrix-vector product or GEMV), where either M or N equals 1, are always memory limited, because their arithmetic intensity is less than 1. Generation speed is therefore limited by memory bandwidth rather than raw compute capacity.
Pope et al. quantified the same asymmetry when they analysed generative inference for large Transformers, reporting a low-batch-size generation latency of 29 ms per token alongside 76% model FLOPS utilisation during large-batch processing of input tokens, and analysing the two regimes separately. For long-context inference requirements, colocating both phases means your compute-heavy Tensor Cores sit underutilized during generation, while active decodes are throttled whenever a large prompt enters the scheduler.
| Inference Phase | Core Kernel Operation | Hardware Bottleneck | Typical Arithmetic Intensity (FLOP/B) | Scaling Driver |
|---|---|---|---|---|
| Prefill | General Matrix Multiplication (GEMM) | Compute (Tensor Cores / TFLOPS) | > 1000 | Prompt token arrival rate |
| Decode | General Matrix-Vector Multiplication (GEMV) | Memory Bandwidth (TB/s HBM) | > 1 | Concurrent active sequences |
The Stuttering Stream: Symptoms of Colocation
In conventional continuous batching architectures, incoming requests enter the active iteration batch alongside established generation steps. When a request with a 64k or 128k prompt arrives, the scheduler executes a massive prefill forward pass. For that iteration, all ongoing generation requests in the batch must wait for the prefill matrix operations to clear the execution queue.
End users do not perceive this as an elevated average response time. Instead, they experience a stuttering stream: an output stream that generates smoothly, stalls completely for 400 to 1200 milliseconds while a new prompt prefills, and then resumes emitting tokens. This phenomenon creates severe tail latency spikes in Inter-Token Latency (ITL) and Time Per Output Token (TPOT).
The Limitations of Chunked Prefill
Serving engines attempt to mitigate this interference using chunked prefill, which breaks large prompts into smaller token chunks (such as 512 or 2048 tokens) and piggybacks them into iterations with active decode tokens. In vLLM production deployments, chunked prefill is enabled by default in the V1 engine, where the max_num_batched_tokens parameter serves as the primary dial between Time to First Token (TTFT) and ITL.
However, chunked prefill does not eliminate phase interference; it merely spreads the latency cost across multiple iterations. Setting chunk sizes small protects ITL but severely prolongs TTFT and introduces extra memory loading overhead, as previous KV chunks must be repeatedly re-read from HBM into cache. vLLM's own disaggregated prefilling documentation concedes that chunked prefill with a proper chunk size also can achieve the same goal, but that in practice it is hard to figure out the correct chunk size value. Under heavy concurrency and sustained long-context traffic, chunked prefill becomes difficult to tune reliably.
Decoupling SLOs with Disaggregated Serving
Prefill/decode (PD) disaggregation resolves this architectural tension by physically splitting the inference pipeline across specialized worker pools. Dedicated prefill instances handle prompt ingestion, generate the initial KV cache, and emit the first token. The accumulated KV tensors are then transmitted to a dedicated decode pool responsible solely for autoregressive generation.
Physical separation decouples Time to First Token from Inter-Token Latency. Infrastructure operators can scale prefill instances according to prompt-token ingress volume and target TTFT SLOs, while independently provisioning decode capacity based on concurrent stream count and target token-generation rates. A sudden burst of long-context documents on the prefill tier can no longer starve or interrupt ongoing decodes.
Throughput versus Goodput Optimization
A frequent misconception among infrastructure leads is that disaggregation intrinsically boosts raw system throughput. vLLM's disaggregated prefilling page says the opposite, in capitals: "Disaggregated prefill DOES NOT improve throughput." The two reasons it does name are tuning TTFT and inter-token latency separately and controlling tail ITL, and the feature is labelled experimental.
The metric that disaggregation directly optimizes is goodput. In the DistServe paper, Zhong et al. define per-GPU goodput as the maximum request rate that can be served adhering to the SLO attainment goal (say, 90%) for each GPU provisioned, and report serving 7.4x more requests or meeting a 12.6x tighter SLO than state-of-the-art colocated systems while staying within latency constraints for more than 90% of requests. By eliminating cross-phase resource contention, disaggregation allows each GPU pool to operate closer to saturation without violating latency targets.
- TTFT Isolation: Prefill pools scale strictly against prompt token arrival rates without being constrained by active generation memory ceilings.
- Tail ITL Stability: Decode instances run continuous autoregressive batches with near-zero jitter, eliminating multi-hundred-millisecond pauses.
- Independent Tensor Parallelism: High tensor parallelism (TP) can be assigned to the compute-heavy prefill phase while decode runs with wider pipeline or sequence parallelism.
- Vendor Benchmark Context: Where a vendor does report a throughput gain, it is measured against its own colocated baseline on specialized hardware. NVIDIA reports up to 30x more requests served for DeepSeek-R1 671B on GB200 NVL72 with Dynamo disaggregated serving at FP4 with 32K input and 8K output tokens, versus in-flight batching on the same platform, not against a tuned single-stage server.
The Cost of Separation: KV Cache Transfer
Disaggregation introduces an inescapable operational overhead: every request processed by a prefill worker must transfer its full Key-Value cache across the network fabric to a designated decode worker before generation can begin. The feasibility of disaggregation depends entirely on whether this transfer penalty is smaller than the latency saved by eliminating colocation interference.
For a large model using Grouped-Query Attention (GQA), a long prompt generates gigabytes of KV cache tensors that have to move before the first decode step can run. Transferring that volume across commodity Ethernet introduces transport latency that can wipe out the TTFT improvement entirely and leave the decode worker waiting.
Transport Machinery: NVLink, RDMA, and NIXL
To render network overhead negligible, disaggregated serving requires high-throughput transport mechanisms. Within a single multi-GPU node, intra-node NVLink provides up to 900 GB/s to 1.8 TB/s of bidirectional bandwidth, allowing near-instantaneous KV transfer. Across multi-node clusters, disaggregation mandates Remote Direct Memory Access (RDMA) over InfiniBand or RoCE v2 networks operating at 400 Gb/s to 800 Gb/s.
Modern orchestration systems implement specialized memory managers to coordinate these transfers. NVIDIA Dynamo pairs the KV Block Manager, a unified memory layer that allocates and remotely shares KV blocks for engines such as vLLM, SGLang and TensorRT-LLM, with the NVIDIA Inference Transfer Library (NIXL), a vendor-agnostic point-to-point data movement library that supports RDMA and GPU-initiated networking behind a non-blocking API. Similarly, vLLM's disaggregated prefilling page documents connector backends including NixlConnector, LMCacheConnectorV1 and MooncakeConnector, and states that vLLM relies on third-party connectors for production-level disaggregated prefilling. On slow commodity networks lacking RDMA, disaggregated serving will consistently perform worse than a well-tuned monolithic instance.
Asymmetric Capacity Planning
The primary economic benefit of prefill/decode disaggregation is the ability to match specific GPU hardware profiles directly to phase workloads, avoiding the costly practice of overprovisioning top-tier accelerators across the entire cluster.
In a monolithic serving architecture, teams frequently deploy uniform clusters of high-cost accelerators (such as NVIDIA H100 or H200 GPUs) simply to secure sufficient compute for prompt ingestion while simultaneously holding massive KV caches for generation. This configuration bleeds capital because memory-bound decode iterations leave expensive 8-bit floating point (FP8) Tensor Cores idle for the vast majority of their operational cycles.
Hardware Mapping Strategies
With a disaggregated architecture, infrastructure teams can map each phase to hardware optimized for its specific bottleneck, reducing the total sovereign inference costs across production fleets:
- Prefill Pool: Provision compute-dense GPUs with maximum TFLOPS per dollar (such as NVIDIA H100 or B200 accelerators). These nodes ingest input tokens at wire speed and operate at high utilization with minimal idle memory reservations.
- Decode Pool: Provision memory-capacity and memory-bandwidth optimized GPUs (such as NVIDIA H200 with 141 GB HBM3e or NVIDIA L40S for smaller quantized models). These nodes maximize concurrent batch capacity without paying for unneeded matrix compute.
- Dynamic Pool Ratios: Workload ratios can be adjusted dynamically. During high-concurrency summarization spikes, prefill-to-decode worker ratios can scale from 1:2 to 1:6 without rebuilding cluster topologies.
Treating inference as two distinct hardware tiers transforms capacity planning from a defensive overprovisioning exercise into a deterministic cost optimization workflow.
When Disaggregation Fails to Pay Off
Disaggregation is not a universal performance upgrade. Adding network coordination and state transfer introduces significant architectural complexity, and running disaggregated pipelines on mismatched workloads will degrade performance and inflate operating costs.
We consistently advise engineering teams to avoid disaggregation when their workload profile exhibits any of the following constraints:
- Short Prompt Contexts: For workloads where input prompts are short, prefill computation completes in a few milliseconds. The latency of serializing and transferring the KV cache will exceed the prefill computation time itself.
- Low Concurrency and Single Replicas: If an application handles low query-per-second (QPS) traffic served on a single GPU or single VM, splitting across multiple instances creates resource fragmentation and guarantees idle GPU waste.
- Commodity Network Fabrics: Deploying disaggregated workers across cloud instances connected only by standard TCP/IP without RDMA or NVLink creates network bottlenecks that severely spike TTFT.
- Generous Latency SLOs: Workloads that do not require strict P99 inter-token latency guarantees (such as offline asynchronous batch pipelines) achieve higher raw hardware efficiency using monolithic continuous batching.
Furthermore, operators must account for the enlarged failure domain. Disaggregation introduces an orchestration proxy, a KV-aware router, network transport daemons, and two distinct autoscaling groups that must balance asymmetric traffic signals. When an issue occurs, debugging state mismatches across distributed worker tiers requires substantially deeper telemetry than troubleshooting a standalone serving engine.
Choosing Infrastructure for a Disaggregated Deployment
Running a disaggregated topology in production means picking infrastructure that matches the transport requirement above. The stack itself is open: runtime engines like vLLM and TensorRT-LLM combined with an orchestration framework such as NVIDIA Dynamo, which reached production-grade 1.0 status with support for SGLang, TensorRT-LLM and vLLM as backends. Dynamo operates above individual inference engines, coordinating multi-node pools, KV Block Management, and KV-aware routing without locking teams into proprietary runtimes.
For engineering teams implementing disaggregated serving topologies, three deployment tiers cover the realistic options:
- On-demand GPU VM: Deploy 1 to 8 GPUs per virtual machine with full NVLink interconnects, per-second billing, and zero egress fees. Ideal for single-node disaggregated experiments where prefill and decode instances communicate across high-speed NVLink memory buses.
- Dedicated Inference: Deploy custom Docker images or Hugging Face model IDs behind private, isolated, auto-scaling endpoints with scale-to-zero capabilities. Served entirely from European data centres in Paris and Finland.
- Large-Scale GPU Cluster: Node-scale capacity interconnected with 400 Gb/s InfiniBand NDR fabrics, managed via Slurm or Kubernetes, for multi-node RDMA disaggregation.
For teams consuming foundation models directly through an OpenAI-compatible API, Lyceum Serverless Inference provides pay-per-token endpoints for leading open-weight architectures. These include 1M-token context models such as MiniMax-M3 ($0.40 per 1M input / $2.00 per 1M output tokens), Kimi-K3 ($3.00 per 1M input / $15.00 per 1M output tokens), GLM-5.2 Instant ($1.50 per 1M input / $4.50 per 1M output tokens), and DeepSeek-V4-Flash ($0.15 per 1M input / $0.30 per 1M output tokens), as well as Qwen3.5-9B featuring a 256K-token window ($0.15 per 1M input / $0.20 per 1M output tokens). Lyceum Serverless Inference is a self-serve offering that carries no formal SLA or service credits.
Whether you are managing your own disaggregated vLLM clusters on InfiniBand compute or calling long-context models via our serverless endpoints, separating prefill and decode transforms long-context serving from a bottleneck into a predictable, scalable architecture.