AI This article was created with the help of AI.
If a GRPO run is running out of memory, this guide accounts for every model copy the algorithm keeps resident and shows which one to attack first. Reinforcement learning workflows frequently crash during initial allocation because teams calculate GPU requirements based solely on the target policy parameters. Group Relative Policy Optimization (GRPO) simplifies reinforcement learning by eliminating the value critic network, but it introduces distinct memory bottlenecks across multiple co-resident models and dynamic rollout buffers.
GRPO keeps more than one model resident
Reinforcement learning with Proximal Policy Optimization (PPO) traditionally requires four concurrent neural networks: the active actor policy, the reference policy, a reward model, and a value critic network that estimates state values for Generalized Advantage Estimation (GAE). GRPO eliminates the critic network entirely by sampling a group of outputs for each input prompt and estimating the baseline from the group scores, a variant of PPO introduced to optimise the memory usage of PPO. This architectural change removes the memory overhead of training a separate value network alongside the actor.
Removing the critic does not mean that only one model resides on your GPUs. An active GRPO pipeline still requires up to three distinct model entities during training iterations: the active policy being optimized, a frozen reference model used to compute the Kullback-Leibler (KL) divergence penalty, and an optional reward model if task verification is not handled by programmatic rule checkers. When estimating GPU memory before a run, calculating capacity for only the trainable policy guarantees a CUDA Out-of-Memory (OOM) fault during the first training step.
Each of these resident models serves a specific mathematical role in the GRPO loss function and exhibits a completely different memory profile. Sizing an infrastructure cluster accurately requires treating each copy as an independent memory consumer rather than multiplying the base parameter count.
Why only the policy carries optimiser state
The active policy is the only model copy in a GRPO run that updates its weights through backpropagation. Because gradients flow through this network, it incurs the full memory overhead of mixed-precision training, gradient storage, and optimizer tracking states. Sizing this component requires accounting for both 16-bit parameter buffers and 32-bit master weight accumulators.
In standard 16-bit mixed-precision training using bfloat16 or float16, the base weights consume 2 bytes per parameter, and the stored gradients consume another 2 bytes per parameter. The dominant memory consumer, however, is the optimizer. The standard AdamW optimizer maintains three 32-bit float values per parameter: a full-precision master copy of the parameter (4 bytes), the first momentum estimate (4 bytes), and the second variance estimate (4 bytes). This adds 12 bytes per parameter in optimizer state alone.
| State Component | Precision | Bytes per Parameter | Memory for 7B Model | Memory for 70B Model |
|---|---|---|---|---|
| Model Weights (BF16) | 16-bit | 2 bytes | 14.0 GB | 140.0 GB |
| Gradients (BF16) | 16-bit | 2 bytes | 14.0 GB | 140.0 GB |
| FP32 Master Weights | 32-bit | 4 bytes | 28.0 GB | 280.0 GB |
| Adam Momentum (FP32) | 32-bit | 4 bytes | 28.0 GB | 280.0 GB |
| Adam Variance (FP32) | 32-bit | 4 bytes | 28.0 GB | 280.0 GB |
| Total Policy Static State | Mixed | 16 bytes | 112.0 GB | 1,120.0 GB |
As demonstrated in the table, the trainable policy demands a baseline of 16 bytes per parameter before accounting for activation memory or communication buffers. The DeepSpeed documentation makes the same point concretely: Adam optimizer states alone consume 18 GB for a 1.5B parameter GPT-2 model, which is 12 bytes per parameter. A 7B parameter policy therefore requires 112 GB of static memory, while a 70B model demands over 1.1 TB across the cluster. Because the policy holds 8 times more memory than a bare 16-bit model file, any optimization applied to the optimizer state yields the largest absolute memory reduction.
Reference and reward copies in inference mode
The reference model and the neural reward model operate strictly in forward-evaluation mode. The reference model evaluates token probabilities to enforce a KL penalty, keeping the policy close to the reference distribution rather than drifting into degenerate outputs or reward hacking. Because its weights remain frozen throughout the entire training process, it requires neither gradient storage nor optimizer states.
A frozen 16-bit reference model consumes exactly 2 bytes per parameter. For a 7B parameter architecture, this represents a fixed footprint of 14 GB, compared to the 112 GB required by the active policy. If your pipeline uses a learned neural reward model of equal parameter scale, it also consumes exactly 2 bytes per parameter in inference mode. In reasoning tasks such as mathematics or code compilation, teams frequently replace the neural reward model with a deterministic reward function or sandbox execution environment, removing the reward model from GPU memory altogether.
- Active Policy: 16 bytes per parameter (Weights + Gradients + FP32 AdamW states)
- Reference Model: 2 bytes per parameter in BF16 (Inference forward pass only)
- Neural Reward Model: 2 bytes per parameter in BF16 (Optional; 0 bytes if using rule verifiers)
- Activation Checkpointing: Adds dynamic activation overhead during the policy backward pass, mitigated via gradient checkpointing
Understanding this asymmetry prevents common sizing mistakes. When a cluster runs out of memory, adding GPUs or sharding policies without distinguishing between the 16-byte active policy and the 2-byte inference copies leads to inefficient distributed layouts.
KV cache for the rollout group on top
Unlike standard supervised fine-tuning where the training dataset is pre-tokenized and static, GRPO generates its own training data on the fly. For every input prompt, the policy samples a group of candidate completions (denoted as group size G) and scores them relative to each other. Generating these completions requires maintaining active Key-Value (KV) cache buffers across the entire rollout batch.
The memory footprint of the KV cache is determined by the number of attention layers, the number of key-value heads, the head dimension, the sequence context length, and the numerical precision. For modern architectures employing Grouped-Query Attention (GQA), the memory per token is significantly lower than standard Multi-Head Attention, but rollout groups quickly multiply this requirement across long reasoning traces, and inefficient management of that cache limits how many sequences can run in a batch. Systems using PagedAttention allocate this memory in dynamic non-contiguous blocks to minimize fragmentation waste.
Generation engines reserve whatever VRAM is left after the model weights are loaded and allocate it to the KV cache per request, so the space available for token generation is whatever the static state does not already occupy. A batch of prompts sampling several rollouts each means the engine must hold many concurrent sequences at once. If the static model weights and optimizer states consume nearly all available VRAM, the generation engine will starve, triggering preemption or runtime crashes before policy backpropagation even begins.
Four levers when it does not fit
When a planned GRPO run exceeds available GPU cluster capacity, you must apply memory reduction techniques systematically. Rather than arbitrarily reducing batch sizes, attack the four primary memory consumers in order of efficiency and computational impact.
| Optimization Lever | Target Component | Typical VRAM Saved | Compute & Throughput Impact | Convergence Impact |
|---|---|---|---|---|
| Eliminate / Offload Reference Model | Reference Model (2 bytes per parameter) | The full reference-model weight footprint | Negligible if offloaded; zero overhead if KL approximated | Zero impact if exact offload; minor if KL omitted |
| Quantize Reference & Reward Models | Inference Copies (FP8 / INT4) | A fraction of the inference-copy footprint, set by target precision | Minimal latency overhead on modern tensor cores | Negligible quality loss on evaluation passes |
| ZeRO-Offload / CPU Sharding | Optimizer States & Gradients | The optimiser-state share of policy memory (12 of 16 bytes per parameter) | PCIe/NVLink transfer overhead reduces step speed | Zero impact on mathematical convergence |
| Reduce Group Size (G) / Sequence Length | Dynamic KV Cache | Linear reduction in cache footprint | Reduces generation compute; smaller advantage sample | Higher gradient variance; may require more steps |
The first lever is addressing the reference model. When high-bandwidth interconnects or host RAM are available, offloading the frozen reference model to CPU memory between forward passes frees its full 2 bytes per parameter of GPU memory, using the same host-offload mechanism DeepSpeed applies to training state. Alternatively, if your reward function incorporates tight length or task constraints, setting the KL divergence coefficient to zero eliminates the reference model requirement completely.
The second lever applies parameter-efficient methods and post-training quantization to inference-only copies. Using Hugging Face PEFT with LoRA allows the base policy to remain frozen while only training low-rank adapter matrices, drastically reducing optimizer state memory. Concurrently, serving the reference and reward models at lower precision shrinks their static footprint, and the generation engine exposes the cache and quantization dtypes as explicit runtime flags.
The third lever is partitioning optimizer states across data-parallel ranks or offloading them to system memory via ZeRO-Stage 1/2/3. Because optimizer states account for 12 of the policy's 16 bytes per parameter, ZeRO partitioning distributes this massive block across available GPUs, allowing larger models to fit comfortably. Finally, tuning generation hyper-parameters like group size G directly controls the KV cache ceiling during rollouts.
Fitting training and rollout together
A common failure mode in reinforcement learning pipelines is treating the training step and the generation phase as independent hardware problems. Sizing a cluster so that model weights and optimizer states occupy 95% of GPU memory appears viable during the backward pass, but it leaves insufficient capacity for the generation engine's KV cache blocks.
When the rollout engine lacks adequate memory headroom, it cannot process group completions in parallel. The engine is forced to serialize rollouts or repeatedly swap KV cache blocks between GPU and host memory, collapsing rollout generation throughput. Because generation typically accounts for 70% to 85% of total wall-clock time in an RL training step, starving the KV cache creates an infrastructure configuration that is technically alive but economically impractical.
- Profile Peak Activation Memory: Account for activation spikes during the policy backward pass with fine-tuning a 70B model architectures before assigning rollout cache limits.
- Set Explicit VRAM Headroom: Configure generation engine memory limits to guarantee that static weights, optimizer states, and activations do not overlap cache pools.
- Colocate or Decouple Engines: For large clusters, consider running generation workers on dedicated inference nodes and streaming trajectories to dedicated training nodes to avoid memory contention.
- Maintain Uniform Interconnect Speeds: Ensure high-bandwidth links across training nodes to prevent ZeRO gather operations from bottlenecking execution steps.
Balancing these memory pools ensures that both the backpropagation step and the high-throughput generation phase execute without resource contention or memory thrashing.
A worked total you can substitute into
To size a cluster before launching a job, compute total resident memory by evaluating each component term explicitly. Consider a full-parameter GRPO run on a 7B parameter base model with an active policy, a frozen reference model, a programmatic rule verifier (0 GB reward model), and an 8-sample rollout group generating up to 4,096 tokens per sequence.
| Memory Component | Formula / Derivation | Memory (GB) |
|---|---|---|
| Policy Weights (BF16) | 7 * 10^9 * 2 bytes | 14.0 GB |
| Policy Gradients (BF16) | 7 * 10^9 * 2 bytes | 14.0 GB |
| Policy AdamW Optimizer States | 7 * 10^9 * 12 bytes (FP32 master, momentum, variance) | 84.0 GB |
| Reference Model (BF16) | 7 * 10^9 * 2 bytes (Inference only) | 14.0 GB |
| Reward Model | Rule-based verifier (0 bytes) | 0.0 GB |
| Activation Checkpointing (Peak) | Estimated activation memory per batch | 12.0 GB |
| Rollout KV Cache (G=8, 4K Context) | Dynamic cache allocation for 8 parallel completions | 18.0 GB |
| CUDA & Framework Overhead | Context buffers, PyTorch caching allocator reserves | 8.0 GB |
| Total Single-Node Requirement | Sum of static, dynamic, and framework memory | 164.0 GB |
In this configuration, total memory across the run equals 164.0 GB. Running this workload without sharding requires at least two 80GB GPUs (such as NVIDIA H100 or A100 nodes) with ZeRO optimizer partitioning or four 80GB GPUs if running full data parallelism without offloading. For larger architectures like 70B models, total static and dynamic state scales past 1.6 TB, requiring multi-node distributed setups with high-speed InfiniBand interconnects.
Managing these multi-model memory profiles across raw infrastructure requires managing complex GPU orchestrations. Teams running RL fine-tuning workloads can eliminate cluster provisioning waste by deploying on Lyceum Serverless Training. Compute the total with every term named, then size the training job against it.