AI This article was created with the help of AI.
Three multipliers before any learning happens
In multi-turn reinforcement learning (RL) for autonomous agents, rollout generation dominates compute spend long before gradient calculation begins. The DeepSeekMath paper introduces Group Relative Policy Optimization (GRPO) as a variant of Proximal Policy Optimization (PPO) that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO, and it does so by scoring a group of sampled outputs per prompt rather than training a separate value network. That structure means the total token computation is governed by three nested multipliers: group size, trajectory turns, and batch prompts. For every policy update, the system must sample multiple independent rollouts across a batch of prompt environments, with each rollout executing several conversational or tool-calling turns.
The disparity between generation work and backward-pass work is stark. A standard supervised fine-tuning step processes input and target tokens in a single parallel forward and backward pass. In contrast, multi-turn RL requires generating thousands of sequential tokens autoregressively across multiple candidate branches just to compute a single reward baseline and policy gradient. When training a model with 7B or 70B parameters, unoptimized rollout loops spend the vast majority of total cluster GPU time simply evaluating environment steps and decoding responses.
Understanding this multiplication makes the scaling bottleneck clear. The actual gradient optimization over the policy network occurs only after all prompt, group, and turn interactions complete. If each interaction naively re-processes token sequences from scratch, compute overhead scales far faster than model parameter size, inflating total cluster requirements and hardware spend.
Every turn re-prefills the whole conversation
The primary driver of unnecessary compute in multi-turn agent RL is duplicated prefill. An agent rollout is inherently stateful: turn 1 begins with a system prompt and an initial task, producing an action or tool call. The environment executes the tool and returns an observation. At turn 2, the agent receives the full history including the system prompt, turn 1 prompt, turn 1 response, and tool observation to generate the next action. If your rollout worker discards the key-value (KV) cache between environment steps, it must re-encode every historical token before generating a single new token.
| Turn Index | Turn Input Tokens | Discarded Cache Prefill Tokens | Persisted Cache Prefill Tokens |
|---|---|---|---|
| Turn 1 | 512 | 512 | 512 |
| Turn 2 | 512 (1,024 cumulative) | 1,024 | 512 |
| Turn 3 | 512 (1,536 cumulative) | 1,536 | 512 |
| Turn 4 | 512 (2,048 cumulative) | 2,048 | 512 |
| Turn T (general) | L tokens per turn | (T * (T + 1) / 2) * L | T * L |
Without cache retention, the number of prefilled tokens across a trajectory scales quadratically with the number of turns. For a trajectory where each turn appends a fixed token length, the naive loop recomputes the entire historical sequence at every step. With cache persistence, the engine computes each token key and value representation exactly once, keeping total prefill proportional to the sequence length. For deep agentic tasks requiring 10 to 20 environment interactions, discarding cache forces the GPU to perform many times more prefill matrix multiplications than mathematically necessary. As the PagedAttention authors put it, the key-value cache memory for each request is huge and grows and shrinks dynamically, so it is worth sizing the tensors yourself using standard approaches to calculating KV cache memory.
Rollouts in a group share their opening prefix
The second major source of redundant computation occurs across the rollout group. GRPO, published as a PPO variant that optimizes the memory usage of PPO, works by sampling a group of outputs from the same prompt and scoring them relative to one another, which is how it avoids the separate value model PPO requires. Every single rollout in that group begins with an identical prefix: system instructions, reasoning guidelines, environment schemas, few-shot demonstrations, and the specific task prompt.
In a standard sequential or un-cached batch implementation, the model computes the attention keys and values for this initial prompt separately for every rollout. For agent workloads where system prompts and tool descriptions consume thousands of tokens, running redundant prefills across all rollouts wastes significant FLOPs before the agent outputs its first token. SGLang's runtime introduces RadixAttention for automatic KV cache reuse across generation calls, and its authors report up to five times higher throughput on workloads including agent, reasoning and multi-turn chat tasks, which is exactly the shape of a rollout group branching from one shared prompt.
Putting a serving engine in the training loop
Most ML training frameworks are designed around static tensor allocations and forward-backward passes. They are not engineered to manage dynamic KV cache memory across long-lived, branching generation sessions. When executing agent rollouts inside a pure training loop, calling standard generation functions typically allocates contiguous VRAM buffers for maximum sequence lengths and deallocates the cache the moment generation finishes.
- Contiguous vs paged memory: training loops allocate static, contiguous buffers, whereas PagedAttention is an attention algorithm inspired by the classic idea of virtual memory and paging in operating systems, which lets vLLM hold memory waste to under 4 percent of the KV cache.
- Ephemeral vs persistent KV caching: standard generation routines flush tensor activations after each turn, while vLLM keeps cached blocks alive and evicts the least recently used block from the head of its free queue when space is needed.
- Static batching vs dynamic continuous batching: serving runtimes dynamically interleave prefill and decode phases across heterogeneous rollout lengths to maximise Tensor Core saturation.
- Tree-based prefix sharing: SGLang's runtime retains the KV cache for prompts and generation results in a radix tree, enabling prefix search, insertion and eviction across the multiple generation calls that make up one language model program instead of recomputing each branch.
By decoupling rollout generation from the gradient update and routing generation requests through an inference engine, the training architecture gains operating-system-style memory management. vLLM caches the KV blocks of processed requests and reuses them whenever a new request arrives with the same prefix, an optimisation its own design documentation calls almost a free lunch because it does not change model outputs. The training policy weights are periodically synchronized to the inference engine, which executes branching rollouts with optimal memory reuse.
Reusing cache across turns and across a group
Implementing KV cache reuse across both turns and groups transforms the computational complexity of agentic RL. When an inference engine manages rollouts, Turn 1 of Rollout 1 prefills the shared prompt and writes the key-value tensors into paged cache blocks; because those blocks need not be contiguous, different sequences can map their logical blocks to the same physical block, and vLLM tracks reference counts on the physical blocks to keep that sharing safe. When subsequent rollouts in the same group initialize, the scheduler hashes the prompt tokens, looks up the already computed blocks, and touches them by increasing their reference count so they are not evicted rather than recomputing them, so generation starts from the cached state.
- Initial Group Prefill: The engine computes KV tensors for the common system prompt and task instructions once, assigning them immutable page identifiers in the radix tree.
- Parallel Generation Branching: Group rollout workers generate independent actions from the cached root prefix without duplicating initial attention computation.
- Stateful Turn Resumption: When an environment returns tool observations for the next turn, the engine appends only the new observation tokens to the existing trajectory branch, prefilling incremental tokens instead of the entire conversation history.
- Branch Eviction: Once all trajectories in a group complete and rewards are collected, the engine evicts child branch pages using a least-recently-used policy while retaining hot system prefixes.
This architecture shifts generation scaling from quadratic compute to linear compute with respect to trajectory length. Across a full training step involving multiple prompt groups and several rollouts per group, eliminating duplicated prefix and turn computation substantially reduces rollout time, allowing GPU clusters to focus on policy convergence rather than redundant matrix operations.
Where memory bounds how much you can keep
While KV cache reuse eliminates redundant prefill compute, it introduces a physical constraint: VRAM capacity. Maintaining active KV caches for multiple concurrent long-context agent trajectories requires substantial GPU memory. For a 16-bit precision model using Grouped Query Attention (GQA), each token's KV cache requires memory proportional to the number of layers, head dimensions, and key-value heads. Across long context windows and high concurrency, cache memory can exceed the memory required for model weights alone.
When training and rollout generation share the same physical GPUs, available memory is further constrained by optimizer states and gradient buffers. Adopting partitioned distributed training strategies like DeepSpeed ZeRO frees per-GPU memory by sharding optimizer states across data-parallel ranks. Freeing this static memory footprint provides the headroom needed to expand the paged KV cache pool without triggering CUDA out-of-memory errors. How much of that headroom the engine actually claims is a serving-side setting rather than a training-framework one: in vLLM the cache and memory behaviour is controlled through engine arguments passed to the LLM class for offline inference or to vllm serve for online serving, so size it deliberately alongside estimating GPU memory before a run.
In modular infrastructure setups, teams often partition their hardware pool into dedicated generation workers and dedicated policy trainer workers. This isolation prevents memory contention between deep RL backward passes and high-concurrency rollout caching, ensuring predictable throughput during massive trajectory exploration.
Measuring prefill separately from decode
To determine whether your training pipeline suffers from duplicated prefill, you must instrument your rollout telemetry to track prefill and decode metrics independently. Aggregating total rollout time obscures the bottleneck: a slow rollout could stem from high decode token counts or from re-encoding massive context windows on every turn. Tracking Time To First Token (TTFT) and Time Per Output Token (TPOT) across every turn exposes exactly where compute cycles are spent.
- Prefill Token Volume: Log the exact number of prompt tokens processed per turn to identify whether input length scales cumulatively or incrementally.
- TTFT per Turn: Measure latency from environment observation submission to the first generated action token; sharp increases in turn-by-turn TTFT indicate missing KV cache reuse.
- Cache Hit Rate: Monitor the serving engine prefix cache hit ratio to verify that group prompts and conversation history are matching. The vLLM design docs state that only full blocks are cached, and their worked example shows a block matching just two of its four tokens producing no hit.
- TPOT Latency: Track per-token generation speed during decode to confirm batching efficiency and detect memory-bandwidth throttling.
Once profiling reveals the extent of redundant prefill in your rollout loop, executing these workloads on optimized compute infrastructure becomes essential. Lyceum provides Serverless Training to run containerized, distributed model training workloads on European sovereign GPU infrastructure. With job start times under 60 seconds, per-second billing, and integration with standard container registries, Serverless Training enables engineering teams to deploy decoupled, cache-aware training pipelines on high-performance NVIDIA hardware without managing underlying cluster orchestration.