AI This article was created with the help of AI.

The Compute Divergence Between Training and Generation

Reinforcement learning post-training loops such as GRPO, PPO, and online DPO force a collision between two fundamentally incompatible workload profiles. On one side sits rollout generation, which runs autoregressive decoding on an inference engine such as vLLM or SGLang. Rollout is strictly memory-bandwidth-bound. Because each generated token requires reading the entire parameter set and managing dynamically allocated KV cache memory, arithmetic intensity remains low. To sustain throughput without drowning in inter-GPU communication latency, rollout workers typically favor modest tensor parallelism degrees (such as TP=1 or TP=2) paired with wide data parallelism across many independent inference replicas.

On the other side sits the training worker, which runs forward activations, loss evaluation, backward gradient passes, and optimizer updates. Training is compute-bound, demanding massive tensor-core FLOPS and high arithmetic intensity. Its memory floor is dictated not by KV cache capacity, but by model weights, gradient shards, optimizer states (such as AdamW first and second moments), and activation buffers. When both roles are co-located on identical silicon without structural separation, the hardware is inevitably misallocated: expensive training FLOPS sit stalled during autoregressive decode steps, or high-bandwidth memory capacity sits starved during dense matrix multiplications.

Workload DimensionRollout / Generation RoleTrainer / Policy Update Role
Primary Hardware BottleneckMemory bandwidth (HBM) and KV cache capacityCompute throughput (TFLOPS) and interconnect bandwidth
Arithmetic IntensityLow (autoregressive token-by-token decode)High (dense GEMM forward and backward passes)
Memory Footprint DriverKV cache blocks and dynamic sequence contextsModel weights, gradients, optimizer states, activations
Optimal Parallelism StrategyLow Tensor Parallelism (TP=1 or TP=2) + Wide Data ParallelismFSDP / ZeRO-3 sharding + moderate Tensor Parallelism
Interconnect SensitivityLow during decode; bursty during weight broadcastsContinuous all-reduce / reduce-scatter collective communications

Understanding this divergence is the prerequisite for dividing your GPU budget. Determining how to split GPUs between trainer and rollout in RL post-training is not a matter of subjective preference or framework defaults. It is an empirical systems optimization problem governed by step-level execution phases, synchronization barriers, and hardware parallelism constraints.

Establishing the Baseline: Colocated vs. Disaggregated Regimes

Before tuning an allocation ratio, you must define the execution regime of your cluster architecture. In a colocated setup, there is no physical GPU split to calculate. Every GPU in the cluster participates in generation, and subsequently every GPU transitions to executing the forward and backward training passes. In this time-multiplexed configuration, your only optimization levers are the memory allocation fraction between the inference engine and the training framework, and the speed at which you can tear down or swap the KV cache before reloading optimizer states.

A true GPU allocation ratio exists exclusively in a disaggregated architecture, where the rollout pool and the training pool remain permanently resident on physically isolated nodes. Disaggregation eliminates the expensive runtime overhead of repeatedly allocating and evicting hundreds of gigabytes of optimizer states and KV cache blocks on the same physical devices. However, this isolation introduces a hard dependency barrier: training workers cannot compute loss or policy gradients until rollout workers finish generating their trajectory batches, and rollout engines cannot begin the next generation cycle until fresh model weights are broadcast from the trainer.

  • Colocated Regime: Dynamic time-sharing on a single unified GPU pool. Zero physical ratio tuning; performance depends on KV cache offload speed and in-place memory reclamation.
  • Disaggregated Regime: Static or dynamic partitioning into separate trainer and rollout clusters. Physical GPU ratio directly determines pipeline throughput and idle bubble duration.
  • Network Fabric Dependency: Disaggregated pipelines require high-bandwidth fabrics, such as 400 Gb/s InfiniBand NDR, to transfer multi-gigabyte trajectory batches and broadcast updated model weights across node boundaries without stalling compute.

When orchestrating distributed runs, the efficiency of a disaggregated system depends entirely on matching the aggregate throughput of your generation nodes to the processing speed of your training nodes.

Real-World Starting Points for the Ratio

When establishing an initial configuration for a disaggregated pipeline, engineering teams often look to published cluster baselines. In production reinforcement learning studies, symmetric allocations frequently serve as the foundational starting point. For example, the RollMux cluster scheduling framework evaluated disaggregated RL post-training across a balanced testbed of 328 NVIDIA H20 inference accelerators paired against 328 NVIDIA H800 compute GPUs, demonstrating how disaggregation aligns distinct hardware tiers to specific workload phases.

Similarly, open-source distributed reinforcement learning frameworks publish disaggregated reference configurations, and they are not necessarily symmetric. In verl's one-step-off async recipe, the documented example splits a physical node so that the trainer takes six GPUs per node and the rollout engine takes two, with the accompanying experiment running four generation GPUs against twelve training GPUs on 16 H20 accelerators. The same documentation treats the ratio as a tuning knob rather than a default: its stated ideal is rollout and training phases of comparable duration, diagnosed by monitoring the wait time on the previous generation together with the sequence-length distribution.

Treat any 1:1 split strictly as an uncalibrated hypothesis. A symmetric ratio assumes that the wall-clock time required to generate your target batch of trajectories exactly matches the wall-clock time required to evaluate rewards and compute policy gradients. In practice, factors such as prompt complexity, sequence generation limits, and policy model architectures disrupt this balance immediately. Rather than accepting framework defaults, you must measure your specific workload to derive an empirical ratio.

Instrumenting the Step to Detect Provisioning Imbalance

To determine whether your cluster is starving on compute or memory bandwidth, you must instrument a single training step into four isolated wall-clock duration buckets: generation time (T_rollout), weight synchronization latency (T_sync), trainer compute duration (T_train), and idle bubble time (T_wait). Measuring these granular phase durations reveals the exact bottleneck governing your pipeline.

  1. T_rollout: The wall-clock duration from when prompts are dispatched to the inference engine until the final sequence token in the batch is decoded.
  2. T_sync: The time spent transferring generated tensors to the trainer and broadcasting updated checkpoint weights back to the inference engines.
  3. T_train: The active compute duration consumed by forward activations, logit evaluations, advantage computations, backpropagation, and optimizer steps.
  4. T_wait: The cumulative idle time spent by either cluster waiting at the synchronization barrier for the counterpart role to finish.

Reading this telemetry provides direct operational guidance. If T_rollout dominates the step while the training cluster records high T_wait, your trainer is over-provisioned relative to your generation capacity; you must reallocate GPUs from the trainer to the rollout pool. Conversely, if T_train consumes the bulk of the cycle while inference workers sit idle with filled completion queues, your rollout pool is over-provisioned, and GPUs should be shifted into the training cluster.

A critical diagnostic trap is relying on whole-run GPU utilization averages reported by cluster monitoring tools. Because synchronous on-policy RL alternates between generation and backpropagation, coarse cluster-wide averages will show moderate, apparently balanced utilization across all nodes even when severe starvation exists in one phase. You must sample utilization per role within discrete step phases to isolate the true compute bottleneck.

The Trap of the Long-Tail Sequence Duration

When profiling rollout latency, the most frequent architectural mistake is treating slow generation as an aggregate throughput shortfall. Large language model reasoning rollouts exhibit a heavily skewed, long-tailed response length distribution, and the imbalance is what makes generation dominate the step: in a characterisation of 14B-parameter GRPO training at a 16k maximum response length on veRL, the rollout stage consumed roughly 70% of each training step across math, code and LLM-as-a-Judge tasks, against 21% to 23% for training. Most prompts in a batch terminate early while a small minority run on toward the context limit, so the step is gated by those few trajectories rather than by the median one.

Under synchronous execution barriers, every GPU in the rollout cluster must idle once its assigned short sequences finish, waiting for the few straggler GPUs processing p99 sequences to complete. If you observe that your p50 sequence completion time is a small fraction of your p99 completion time, simply adding more rollout GPUs will yield negligible end-to-end speedups. The additional GPUs will complete their short sequences even faster, only to sit idle longer waiting for the tail.

Metric / Phase FeatureBalanced Rollout ProfileLong-Tail Skewed ProfileArchitectural Implication
p99 to p50 Latency RatioNear parity (uniform response length)Tail runs many times the median (extreme token variance)High ratio indicates straggler stalling, not lack of raw GPU throughput
Rollout SM UtilizationSustained until step endSharp drop once the median sequences completeWasted VRAM and silicon cycles across majority of rollout nodes
Recommended RemediationScale rollout GPU count linearlyImplement streaming trainers or dynamic tail batchingAdding static GPUs to long-tail workloads increases cluster idle waste

To address this imbalance without wasteful over-provisioning, modern RL systems literature explores streaming architectures. The RollPacker framework, evaluated on up to 128 NVIDIA H800 GPUs, introduces a stream trainer that scales down the number of GPUs dedicated to rollout as generation advances and repurposes the freed devices for training, computing gradients on a subset of data-parallel replicas while deferring the weight update until rollout completes so on-policy correctness is preserved. The lesson generalises beyond that one system: the fix for a long tail is reclaiming the idle silicon, not buying more of it.

How Sequence Length and Group Size Shift the Ratio

When tuning a disaggregated cluster, two algorithmic parameters exert the strongest directional pressure on the optimal GPU allocation ratio: maximum sequence response length and the group sampling size (such as the G parameter in GRPO). Understanding how these parameters scale across hardware boundaries prevents systemic under-provisioning when modifying training recipes.

Increasing the maximum generation length shifts the required allocation heavily toward the rollout cluster. In training, token growth increases compute demand linearly, but the operations retain high arithmetic intensity and benefit from full tensor-core utilization. In rollout, however, longer sequences translate to extended sequential autoregressive decoding steps. Because decoding cannot be compressed through standard batching parallelism, generation time scales directly with step count, multiplying memory bandwidth pressure across the KV cache and dramatically extending T_rollout relative to T_train.

Scaling the GRPO group size (for example, generating 16 or 32 sample completions per prompt instead of 8) also moves the optimal GPU ratio toward rollout. For the training worker, processing additional candidate completions per prompt simply increases the effective token batch size, improving GEMM execution efficiency and increasing MFU during the backward pass. For the rollout engine, however, larger group sizes translate to a surge in concurrent active sequences. These simultaneous streams rapidly saturate available GPU VRAM with KV cache allocations, increasing memory contention, triggering request preemptions, and requiring additional rollout replicas to maintain throughput.

Quantising the Ratio to Node and Parallelism Boundaries

In theoretical modeling, an optimal ratio might suggest allocating 5.4 GPUs to rollout for every 2.6 GPUs allocated to training. In physical infrastructure, however, the continuous ratio is strictly quantised by tensor parallelism degrees, pipeline partitions, and bare-metal node boundaries. Every allocation must divide cleanly into whole-GPU allotments that satisfy the minimum tensor parallel degree required to hold each role's memory footprint.

On the trainer side, the memory floor is dictated by model parameters, gradient shards, and optimizer states under FSDP or ZeRO-3. On the rollout side, the floor is determined by model weights plus the KV cache volume required to sustain your concurrency target without out-of-memory errors. The hardware capacity of your silicon dictates the minimum viable partition:

  • NVIDIA L40S (48 GB): Strictly suitable for smaller parameter models or dedicated low-concurrency rollout workers; cannot be deployed in large-scale cluster fabrics.
  • NVIDIA A100 (80 GB) and H100 (80 GB): Standard enterprise baseline; running 70B parameter models requires TP=4 or TP=8 partitions to preserve VRAM for KV cache blocks.
  • NVIDIA H200 (141 GB) and B200 (192 GB): High-capacity HBM architectures that allow large models (such as 70B variants) to run at lower tensor parallelism degrees (e.g., TP=2), unlocking flexible ratio combinations (such as 6:2 or 2:6 on 8-GPU nodes).

For single-node prototyping, an on-demand GPU VM offering 1 to 8 GPUs behind a single NVLink domain is usually the cheapest place to find the reachable set empirically. Within an 8-GPU node boundary, tensor parallel requirements restrict your disaggregated allocations to discrete splits such as 6:2, 4:4, or 2:6, and any implied ratio between those points has to be rounded to one of them before you re-measure.

For distributed production scaling, Lyceum delivers Large-Scale GPU Cluster infrastructure spanning single nodes up to many thousands of GPUs networked over a 400 Gb/s InfiniBand NDR fabric, orchestrated via Slurm or Kubernetes with 28-second provisioning. Available on terms of 3, 6, 12, or 24 months, these dedicated clusters support NVIDIA H100, H200, B200, and B300 hardware hosted in European data centres across Paris and Finland. Service level agreements are established per business contract. Containerised training pipelines can also run as managed Serverless Training jobs starting in under 60 seconds from ECR, GAR, or Docker Hub, giving engineering teams dedicated control over their distributed training topology.