Distributed Runs

9 articles

Articles

26 August 2026

Weight Sync Between Trainer and vLLM: The Hidden Cost in RL Loops

Moving updated policy weights from your trainer to vLLM for rollouts incurs a recurring time penalty that bleeds capital. This guide models the true cost of weight synchronization over disk, NCCL, and delta transfers, explaining how your cluster fabric sets the limit.

24 September 2026

Training MoE Models: Preventing Expert Collapse & Balancing Load

Training a Mixture-of-Experts model introduces a critical bottleneck: if the router favors a small subset of experts, the starved parameters waste memory while overloaded ones halt the distributed run. Here is how to prevent routing collapse and balance MoE loads effectively.

23 September 2026

Training Long-Context Models Without OOM: Ring Attention

When a training run fails on long sequences, the culprit is unsharded activation memory, not model parameters. Standard tensor and pipeline parallelism will not fix it. Context parallelism splits the sequence itself across GPUs, enabling massive contexts without OOMs.

17 September 2026

Multi-Node GRPO Orchestration: Ray, Slurm or Kubernetes

When scaling GRPO to multi-node clusters, Ray is not an alternative to Slurm or Kubernetes; it is the runtime that sits inside them. Discover why RL's co-dependent architecture makes gang scheduling non-negotiable and how to orchestrate your training jobs on Lyceum.

15 September 2026

How Many GPUs for Trainer vs Rollout in RL Post-Training

In a disaggregated reinforcement learning pipeline, colocation is obsolete. Here is how to instrument your RL post-training loop, measure phase-level timings, and properly split your GPU fleet between the compute-bound trainer and the memory-bound rollout engine.

9 September 2026

Agentic RL GPU Cost: Multi-Turn Rollouts and KV Reuse

Agentic reinforcement learning repeats identical prefill calculations across turns and rollouts, compounding GPU hours. Moving a serving engine like vLLM into the training loop stops this waste by reusing the KV cache across the entire group.

22 May 2026

Multi-GPU Tensor Parallelism Setup: Configuration and Optimization Guide

A 70B model needs about 140GB in FP16 and does not fit on one 80GB GPU. Tensor parallelism splits weight matrices across devices, at the cost of four all-reduce collectives per transformer layer in a training step.

16 May 2026

Multi GPU Distributed Training Setup Guide: Frameworks & Infrastructure

Scaling from a single GPU to a multi-node cluster introduces complex communication bottlenecks and fatal memory errors. Learn how to configure DDP, FSDP, and DeepSpeed while optimizing your infrastructure for maximum throughput.

23 February 2026

ZeRO-3 vs FSDP: A Deep Dive into Memory Efficiency for LLMs

Scaling large language models requires moving beyond standard data parallelism to overcome the memory wall. This technical guide compares DeepSpeed ZeRO-3 and PyTorch FSDP to help engineers optimize GPU utilization and eliminate out-of-memory errors.

Your next workload starts here