Distributed Runs
9 articles
Articles
26 August 2026
Weight Sync Between Trainer and vLLM: The Hidden Cost in RL Loops
Moving updated policy weights from your trainer to vLLM for rollouts incurs a recurring time penalty that bleeds capital. This guide models the true cost of weight synchronization over disk, NCCL, and delta transfers, explaining how your cluster fabric sets the limit.
24 September 2026
Training MoE Models: Preventing Expert Collapse & Balancing Load
Training a Mixture-of-Experts model introduces a critical bottleneck: if the router favors a small subset of experts, the starved parameters waste memory while overloaded ones halt the distributed run. Here is how to prevent routing collapse and balance MoE loads effectively.
23 September 2026
Training Long-Context Models Without OOM: Ring Attention
When a training run fails on long sequences, the culprit is unsharded activation memory, not model parameters. Standard tensor and pipeline parallelism will not fix it. Context parallelism splits the sequence itself across GPUs, enabling massive contexts without OOMs.
17 September 2026
Multi-Node GRPO Orchestration: Ray, Slurm or Kubernetes
When scaling GRPO to multi-node clusters, Ray is not an alternative to Slurm or Kubernetes; it is the runtime that sits inside them. Discover why RL's co-dependent architecture makes gang scheduling non-negotiable and how to orchestrate your training jobs on Lyceum.
15 September 2026
How Many GPUs for Trainer vs Rollout in RL Post-Training
In a disaggregated reinforcement learning pipeline, colocation is obsolete. Here is how to instrument your RL post-training loop, measure phase-level timings, and properly split your GPU fleet between the compute-bound trainer and the memory-bound rollout engine.
9 September 2026
Agentic RL GPU Cost: Multi-Turn Rollouts and KV Reuse
Agentic reinforcement learning repeats identical prefill calculations across turns and rollouts, compounding GPU hours. Moving a serving engine like vLLM into the training loop stops this waste by reusing the KV cache across the entire group.
22 May 2026
Multi-GPU Tensor Parallelism Setup: Configuration and Optimization Guide
A 70B model needs about 140GB in FP16 and does not fit on one 80GB GPU. Tensor parallelism splits weight matrices across devices, at the cost of four all-reduce collectives per transformer layer in a training step.
16 May 2026
Multi GPU Distributed Training Setup Guide: Frameworks & Infrastructure
Scaling from a single GPU to a multi-node cluster introduces complex communication bottlenecks and fatal memory errors. Learn how to configure DDP, FSDP, and DeepSpeed while optimizing your infrastructure for maximum throughput.
23 February 2026
ZeRO-3 vs FSDP: A Deep Dive into Memory Efficiency for LLMs
Scaling large language models requires moving beyond standard data parallelism to overcome the memory wall. This technical guide compares DeepSpeed ZeRO-3 and PyTorch FSDP to help engineers optimize GPU utilization and eliminate out-of-memory errors.