Fine-Tuning

5 articles

Articles

22 September 2026

Gradient Accumulation Loss Bug: Why Effective Batch Is Wrong

Gradient accumulation is assumed to be mathematically identical to full-batch training, but the standard implementation normalizes over the wrong denominator. Here is why the mean-of-means error skews weights, and how to verify if your fine-tuning setup is affected.

14 September 2026

GRPO vs DPO vs RLHF: Compute Cost and GPU Footprint Compared

Direct Preference Optimization (DPO), RLHF, and GRPO scale their compute footprints differently based on how many models they keep resident. Here is the definitive comparison of resident model headcounts, generation wall-clock time, and memory bandwidth constraints.

10 September 2026

GRPO VRAM & GPU Sizing: Policy, Reference and Reward

If a GRPO run is running out of memory, this guide accounts for every model copy the algorithm keeps resident. Sizing GPUs correctly requires accounting for the policy's optimiser state, inference copies, and the rollout cache before you start.

20 May 2026

LoRA vs Full Fine-Tuning Memory Cost: VRAM Math

You have a 24GB GPU and an 8B model. The math says it should fit, but your training script crashes with an OOM error before the first epoch. We break down the exact VRAM requirements for full fine-tuning versus LoRA.

18 May 2026

FP8 Training on H100: Benchmarks and Memory Savings

Training a 70-billion parameter model in BF16 requires hundreds of gigabytes of GPU memory. Shifting to FP8 precision on NVIDIA H100s halves the bytes per element for the tensors actually held in FP8, master weights and optimizer states stay in higher precision, and NVIDIA's NeMo measurements show 1.30x throughput on Llama 3 8B and 1.43x on Llama 3 70B versus BF16.

Your next workload starts here