Training Infrastructure
Training and fine-tuning infrastructure: OOM prevention, VRAM estimation, memory profiling, multi-GPU and multi-node scaling, large datasets.
17 articles
Articles
23 February 2026
Maximizing VRAM: Gradient Checkpointing Memory Savings Guide
26 August 2026
Weight Sync Between Trainer and vLLM: The Hidden Cost in RL Loops
Moving updated policy weights from your trainer to vLLM for rollouts incurs a recurring time penalty that bleeds capital. This guide models the true cost of weight synchronization over disk, NCCL, and delta transfers, explaining how your cluster fabric sets the limit.
24 September 2026
Training MoE Models: Preventing Expert Collapse & Balancing Load
Training a Mixture-of-Experts model introduces a critical bottleneck: if the router favors a small subset of experts, the starved parameters waste memory while overloaded ones halt the distributed run. Here is how to prevent routing collapse and balance MoE loads effectively.
22 September 2026
Gradient Accumulation Loss Bug: Why Effective Batch Is Wrong
Gradient accumulation is assumed to be mathematically identical to full-batch training, but the standard implementation normalizes over the wrong denominator. Here is why the mean-of-means error skews weights, and how to verify if your fine-tuning setup is affected.
23 September 2026
Training Long-Context Models Without OOM: Ring Attention
When a training run fails on long sequences, the culprit is unsharded activation memory, not model parameters. Standard tensor and pipeline parallelism will not fix it. Context parallelism splits the sequence itself across GPUs, enabling massive contexts without OOMs.
21 September 2026
Sequence Packing: Reclaiming GPU Hours in Long-Context SFT
Padding waste silently inflates long-context SFT costs by processing empty tokens. This guide explains how to calculate token occupancy, migrate from length-grouped batching to true sequence packing, and prevent the three silent correctness bugs that destroy model quality.
16 September 2026
Checkpoint Resume Loss Spikes: Fixing Optimizer & LR State
When a long training run is interrupted, resuming from a checkpoint often triggers a massive loss spike. The most common culprits are missing optimizer moment buffers, reset learning rate schedulers, or mismapped FSDP shards - here is the triage order and recovery checklist.
17 September 2026
Multi-Node GRPO Orchestration: Ray, Slurm or Kubernetes
When scaling GRPO to multi-node clusters, Ray is not an alternative to Slurm or Kubernetes; it is the runtime that sits inside them. Discover why RL's co-dependent architecture makes gang scheduling non-negotiable and how to orchestrate your training jobs on Lyceum.
15 September 2026
How Many GPUs for Trainer vs Rollout in RL Post-Training
In a disaggregated reinforcement learning pipeline, colocation is obsolete. Here is how to instrument your RL post-training loop, measure phase-level timings, and properly split your GPU fleet between the compute-bound trainer and the memory-bound rollout engine.
14 September 2026
GRPO vs DPO vs RLHF: Compute Cost and GPU Footprint Compared
Direct Preference Optimization (DPO), RLHF, and GRPO scale their compute footprints differently based on how many models they keep resident. Here is the definitive comparison of resident model headcounts, generation wall-clock time, and memory bandwidth constraints.
10 September 2026
GRPO VRAM & GPU Sizing: Policy, Reference and Reward
If a GRPO run is running out of memory, this guide accounts for every model copy the algorithm keeps resident. Sizing GPUs correctly requires accounting for the policy's optimiser state, inference copies, and the rollout cache before you start.
9 September 2026
Agentic RL GPU Cost: Multi-Turn Rollouts and KV Reuse
Agentic reinforcement learning repeats identical prefill calculations across turns and rollouts, compounding GPU hours. Moving a serving engine like vLLM into the training loop stops this waste by reusing the KV cache across the entire group.
22 May 2026
Multi-GPU Tensor Parallelism Setup: Configuration and Optimization Guide
A 70B model needs about 140GB in FP16 and does not fit on one 80GB GPU. Tensor parallelism splits weight matrices across devices, at the cost of four all-reduce collectives per transformer layer in a training step.
20 May 2026
LoRA vs Full Fine-Tuning Memory Cost: VRAM Math
You have a 24GB GPU and an 8B model. The math says it should fit, but your training script crashes with an OOM error before the first epoch. We break down the exact VRAM requirements for full fine-tuning versus LoRA.
18 May 2026
FP8 Training on H100: Benchmarks and Memory Savings
Training a 70-billion parameter model in BF16 requires hundreds of gigabytes of GPU memory. Shifting to FP8 precision on NVIDIA H100s halves the bytes per element for the tensors actually held in FP8, master weights and optimizer states stay in higher precision, and NVIDIA's NeMo measurements show 1.30x throughput on Llama 3 8B and 1.43x on Llama 3 70B versus BF16.
16 May 2026
Multi GPU Distributed Training Setup Guide: Frameworks & Infrastructure
Scaling from a single GPU to a multi-node cluster introduces complex communication bottlenecks and fatal memory errors. Learn how to configure DDP, FSDP, and DeepSpeed while optimizing your infrastructure for maximum throughput.
23 February 2026
ZeRO-3 vs FSDP: A Deep Dive into Memory Efficiency for LLMs
Scaling large language models requires moving beyond standard data parallelism to overcome the memory wall. This technical guide compares DeepSpeed ZeRO-3 and PyTorch FSDP to help engineers optimize GPU utilization and eliminate out-of-memory errors.