AI This article was created with the help of AI.
The Resident Model Headcount: PPO, GRPO, and DPO
When planning the compute budget for post-training alignment, the decisive variable is not the loss function formulation or the benchmark score. It is the raw headcount of neural networks that must remain resident in GPU memory simultaneously. Each alignment methodology imposes fundamentally different memory architectures during execution, dictating your minimum GPU count before a single gradient step is calculated.
Proximal Policy Optimization (PPO), the standard reinforcement learning from human feedback (RLHF) framework, maintains four distinct models in memory during active training: the actor policy being trained, a frozen reference model to calculate the Kullback-Leibler (KL) divergence penalty, a reward model scoring completions, and a critic (value) network estimating expected future reward. Group Relative Policy Optimization (GRPO) simplifies this topology by retaining the actor policy and the frozen reference model while eliminating the critic network entirely: the paper that introduced GRPO presents it as a variant of PPO that forgoes the critic model and concurrently optimizes the memory usage of PPO. Instead of fitting a parameterized value function, GRPO computes advantages by normalizing reward scores across a sampled group of completions for the same prompt. Direct Preference Optimization (DPO) strips the system down further to exactly two models: the active policy and the frozen reference copy. Its authors reparameterize the reward model so that the optimal policy can be extracted in closed form, solving the standard RLHF problem with only a simple classification loss and eliminating the need to sample from the language model during fine-tuning. DPO therefore removes both the reward model and the critic, discarding the autoregressive rollout generation phase altogether.
| Alignment Method | Active Policy (Trainable) | Reference Model (Frozen) | Reward Model (Inference) | Critic Network (Trainable) | Total Resident Models |
|---|---|---|---|---|---|
| PPO (RLHF) | 1 (Weights + Gradients + Optimizer) | 1 (Weights + KV Cache) | 1 (Weights + KV Cache) | 1 (Weights + Gradients + Optimizer) | 4 |
| GRPO | 1 (Weights + Gradients + Optimizer) | 1 (Weights + KV Cache) | 1 (Learned Model or Code Verifier) | 0 (Replaced by Group Statistics) | 2 + Scorer |
| DPO | 1 (Weights + Gradients + Optimizer) | 1 (Weights + KV Cache) | 0 (Implicit in Policy Loss) | 0 (None) | 2 |
This baseline headcount establishes the physical floor for your GPU cluster. A post-training pipeline requiring four co-resident networks across tens of billions of parameters forces immediate model sharding, tensor parallelism, and pipeline parallelism across multiple nodes, whereas a two-model setup often fits within a single multi-GPU instance.
Why Trainable Networks Cost Far More Than Frozen Copies
Calculating memory footprints by simply multiplying the resident model headcount by the parameter size produces critical allocation failures. In post-training environments, resident models fall into two fundamentally distinct memory tiers: trainable networks and frozen inference copies. A trainable network carries not only model weights, but also gradients and optimizer states, creating massive memory asymmetry.
Under standard mixed-precision training with the AdamW optimizer, memory consumption is defined by specific accounting conventions rather than a single static number:
- ZeRO Convention (16 bytes per parameter): 2 bytes for FP16/BF16 weights, 2 bytes for FP16/BF16 gradients, 4 bytes for FP32 master weights, 4 bytes for FP32 momentum, and 4 bytes for FP32 variance.
- Hugging Face Default (18 bytes per parameter): 6 bytes across weight representations, 4 bytes for FP32 gradients, and 8 bytes for FP32 optimizer states.
- QLoRA / BF16 Direct (12 bytes per parameter): 2 bytes for BF16 weights, 2 bytes for BF16 gradients, and 8 bytes for optimizer states without an explicit FP32 master copy.
In contrast, a frozen network (such as the reference model or an inference-only reward model) requires only 2 bytes per parameter in BF16/FP16, along with its dynamic Key-Value (KV) cache. Set that 2-byte frozen baseline against the trainable conventions above and the asymmetry is stark: Hugging Face's accounting for a model trained in mixed precision with Adam adds up to 18 bytes per parameter, 6 bytes for the fp16 and fp32 weight copies, 8 bytes for optimizer momentum and variance, and 4 bytes for fp32 gradients, before any activation memory. That gap reveals why PPO is so exceptionally demanding: its critic network is not an inference pass, but a second fully trainable network requiring its own optimizer states and gradients.
| Model Type & Convention | Bytes / Param | 70B Model VRAM (Weights + State) | Memory State Multiplier vs Frozen |
|---|---|---|---|
| Frozen Model (BF16 / FP16 Inference) | 2 B/param | 140 GB | 1.0x (Baseline) |
| Trainable: QLoRA / BF16 Direct | 12 B/param | 840 GB | 6.0x |
| Trainable: ZeRO Mixed-Precision AdamW | 16 B/param | 1,120 GB | 8.0x |
| Trainable: Hugging Face Mixed-Precision | 18 B/param | 1,260 GB | 9.0x |
For an engineering team fine-tuning a 70B parameter model, a single trainable copy therefore spans hundreds of gigabytes of raw state, several times the model's inference weights, before accounting for activation buffers or sequence KV cache. When sizing infrastructure, always budget additively (Weights + KV Cache + Activations + Framework Overhead) rather than using arbitrary percentage buffers, and verify whether hardware vendors quote decimal gigabytes (GB) or binary gibibytes (GiB). For deep dive formulas and parameter breakdowns, consult our guide on GPU memory estimation alongside our analysis of LoRA vs full fine-tuning.
The Hidden Pre-Training Cost of Reward Models
Discussions of post-training economics often examine only the active alignment run, ignoring the substantial compute expended beforehand. For standard RLHF, the alignment stage cannot begin until a dedicated reward model has been trained, evaluated, and calibrated on human preference datasets.
Training an enterprise-grade reward model is an independent, resource-intensive machine learning project. It requires curating thousands of prompt-completion pairs, executing full-parameter supervised fine-tuning runs, and mitigating reward hacking risks. If the reward model exhibits calibration drift or low accuracy, the downstream PPO policy rapidly exploits those weaknesses, necessitating complete retries of both reward modeling and policy optimization.
- RLHF (PPO): Requires a separate, upfront supervised fine-tuning run to train a standalone reward model on paired comparisons, followed by co-hosting the trained reward model during policy updates.
- DPO: Completely eliminates the reward model pre-training phase. By reparameterizing the reward so that the optimal policy can be extracted in closed form, DPO optimizes the actor model directly on static preference pairs with a simple classification loss.
- GRPO: Operates flexibly across two regimes. In open-ended domains it scores completions with a reward model, while in verifiable domains such as mathematics and competitive coding the reward can come from checking the answer itself, which is how DeepSeek reports incentivizing reasoning through pure reinforcement learning without human-labeled reasoning trajectories.
In verifiable reasoning tasks, GRPO completely removes human annotation costs and reward model pre-training overhead, generating accurate reinforcement signals on the fly via sandboxed code execution.
Generation vs. Gradient: Where the Wall-Clock Time Goes
The execution profiles of these three alignment methodologies differ fundamentally in how they distribute wall-clock compute time between forward-backward gradient computation and autoregressive rollout generation.
DPO execution resembles traditional Supervised Fine-Tuning (SFT). Because it processes a pre-collected, static dataset of chosen and rejected response pairs, DPO requires no online text generation during training. The compute workload consists strictly of batched forward passes through the policy and reference models, followed by a single backward pass through the policy. The execution time is stable, predictable, and bound almost entirely by tensor core matrix multiplication FLOPs.
In sharp contrast, GRPO and PPO are dominated by on-policy token generation. In GRPO, every optimization step begins by sampling a group of completions for each input prompt. Total wall-clock time is governed by the rollout generation phase according to a specific scaling relationship:
- Rollout Volume = Prompts per Batch * Group Size (G) * Completion Sequence Length
- Active Rollout Time = Rollout Volume / Aggregate System Decode Throughput (tokens/sec)
- Gradient Step Time = Backward Pass Duration (minimal relative to total rollout duration)
When reporting or calculating decode throughput, infrastructure engineers must state the operational parameters explicitly: model parameter size, numerical precision (e.g., FP8 vs BF16), GPU SKU, tensor/pipeline parallelism split, concurrency, and whether the metric represents aggregate cluster tokens per second or single-stream per-user latency. In multi-GPU environments with continuous batching, aggregate system throughput is a large multiple of the single-stream figure, so quoting one where the other is meant will distort a rollout time estimate by orders of magnitude.
Hardware Impact: Memory Bandwidth and Decode Throughput
Because the generation phase dominates GRPO and RLHF wall-clock time, hardware selection must be guided by memory bandwidth rather than theoretical compute FLOPs. Autoregressive token decoding is strictly memory-bandwidth bound: for every generated token, the entire model weight tensor must be read from High Bandwidth Memory (HBM) to the processor compute cores.
This physical bottleneck is illustrated clearly when comparing Hopper-class architectures. The NVIDIA H200 carries the same tensor core throughput as the H100, with NVIDIA publishing 1,979 TFLOPS of FP16 and BFLOAT16 tensor core performance with sparsity for the H200 SXM alongside 141 GB of HBM3e at 4.8 TB/s. The H100 SXM offers 80 GB of HBM3 at 3.35 TB/s, so identical compute is fed by materially less bandwidth: 4.8 divided by 3.35 gives roughly a 1.43 ratio in favour of the H200.
| GPU Architecture | VRAM Capacity | Memory Bandwidth | Bandwidth vs H100 SXM | Frozen 70B BF16 Copies (140 GB) per 8-GPU Node |
|---|---|---|---|---|
| NVIDIA L40S | 48 GB GDDR6 | 864 GB/s | 0.26x (864 / 3,350) | 2 (8 x 48 GB = 384 GB aggregate / 140 GB) |
| NVIDIA A100 SXM | 80 GB HBM2e | 2,039 GB/s | 0.61x (2,039 / 3,350) | 4 (8 x 80 GB = 640 GB aggregate / 140 GB) |
| NVIDIA H100 SXM | 80 GB HBM3 | 3,350 GB/s | 1.00x (baseline) | 4 (8 x 80 GB = 640 GB aggregate / 140 GB) |
| NVIDIA H200 SXM | 141 GB HBM3e | 4,800 GB/s | 1.43x (4,800 / 3,350) | 8 (8 x 141 GB = 1,128 GB aggregate / 140 GB) |
That bandwidth advantage translates directly into higher token generation throughput during the GRPO rollout loop, substantially shortening total wall-clock duration. However, engineers must treat theoretical bandwidth as an upper ceiling rather than a prediction: real decode engines achieve only a fraction of peak memory bandwidth, a fraction captured by the model bandwidth utilization metric, and it tends to fall as peak bandwidth rises. Furthermore, alternating between generation and gradient steps leads to GPU compute underutilization unless rollout engines and training processes are carefully co-scheduled across nodes.
Two Structural Levers to Reduce GPU Footprint
Beyond selecting between algorithms, infrastructure leads have access to two structural architectural levers that significantly lower the hardware requirements of post-training pipelines.
The first lever applies directly to DPO: reference log-probability caching. Because DPO optimizes over a static preference dataset, the reference model is entirely frozen and its outputs do not change between training epochs. Engineers can run a single offline inference pass across the dataset to compute and store the reference log-probabilities for all chosen and rejected tokens on disk. During active training, the reference model is completely eliminated from GPU memory, reducing the DPO footprint from two resident models down to a single active policy.
The second lever applies to online reinforcement learning (GRPO and PPO): modifying or omitting the reference policy KL divergence penalty. In traditional formulations, evaluating the per-token KL divergence term forces an active forward pass through the reference model for every sampled rollout batch. In verifiable reasoning tasks, several modern GRPO implementations rely on strict rule-based verification or clipping bounds, allowing the frozen reference model to be offloaded to host memory or entirely bypassed during gradient steps, eliminating its co-resident VRAM allocation.
- DPO Logprob Caching: Precompute reference probabilities offline, so the run holds only the active policy resident instead of policy plus reference.
- KL Term Omission in GRPO: Replace active reference forward passes with bounded rule penalties, freeing VRAM otherwise occupied by reference model weights and KV caches.
The Honest Trade-Off: Offline Efficiency vs. On-Policy Exploration
Choosing an alignment strategy requires balancing computational cost against model capability. DPO provides an exceptionally cost-effective, operationally stable workflow by operating as an offline method over fixed preference data. However, because DPO cannot sample beyond its static dataset, it cannot autonomously explore new reasoning trajectories or discover self-correcting logic.
GRPO and RLHF incur the compute cost of iterative on-policy sampling in exchange for continuous exploration. In complex tasks like multi-step mathematics, structured coding, and formal verification, on-policy exploration lets models develop reasoning behaviours that were not present in the demonstration data: DeepSeek reports that pure reinforcement learning on a base model facilitates the emergent development of advanced reasoning patterns such as self-reflection, verification, and dynamic strategy adaptation, without human-labeled reasoning trajectories.
For ML teams executing post-training workloads in Europe, Lyceum provides infrastructure tailored for both lightweight fine-tuning and massive reinforcement learning clusters:
- Serverless Training: Submit training or fine-tuning jobs that containerize automatically and launch in under 60 seconds with S3-compatible storage and direct registry integration.
- On-demand GPU VM: Direct SSH access across 1 to 8 GPUs per instance with sub-20-second provisioning and per-second billing. Multi-GPU instances on SXM architectures utilize NVLink, while L40S instances utilize PCIe Gen4 (64 GB/s).
- Large-Scale GPU Cluster: Dedicated clusters scaling up to multi-node configurations interconnected via 400 Gb/s InfiniBand NDR, orchestrated through Slurm or Kubernetes.
All compute capacity is deployed across Lyceum European data centres in Paris and Finland, maintaining full EU data residency and compliance standards. Hardware options range from NVIDIA L40S (48 GB) and A100 (80 GB) up to NVIDIA H100 (80 GB), H200 (141 GB), and B200, whose catalogue capacity is 192 GB of architectural HBM3e while NVIDIA publishes 180 GB usable per GPU on the shipping HGX and DGX B200 systems, so any fits or does-not-fit reasoning should run against the 180 GB figure. For current deployment pricing and node reservations, review the live options on our pricing page.