Diagnosing the Sequence-Length OOM: Activations vs. States

When a training run executes cleanly at 4K tokens but throws a CUDA Out-of-Memory error the moment you extend the context to 32K or 64K, the model itself has not changed. The static memory footprint consisting of model weights, gradients, and optimizer states remains constant regardless of sequence length. Instead, extended sequences directly inflate dynamic activation memory, which must be retained in VRAM for gradient computation during the backward pass.

In traditional transformer architectures, activation memory per layer follows the theoretical upper-bound derivation established by Korthikanti et al. in their 2022 paper on reducing activation recomputation:

  • Activation Memory per Layer = s * b * h * (34 + 5 * a * s / h) bytes for a GPT-style Multi-Head Attention layer with a 4h GELU MLP at 16-bit precision, where s represents sequence length, b denotes microbatch size, h is hidden dimension, and a is the number of attention heads.
  • The term 5 * a * s / h represents the materialization of the full s-by-s attention score matrix (Query-Key scores and softmax outputs).
  • Modern implementations utilizing FlashAttention kernels avoid materializing this quadratic s-by-s matrix in high-bandwidth memory (HBM) entirely, using tiling to compute attention blockwise in on-chip SRAM so that memory use is linear in sequence length.

FlashAttention reduces the practical activation scaling from quadratic to roughly linear with sequence length. However, at extreme sequence lengths (32K to 1M tokens), even linear growth causes activations to consume dozens of gigabytes per GPU per layer. When a run crashes at 64K tokens, the failure is almost entirely an activation footprint issue that static sharding cannot resolve.

Why the Parallelism You Already Have Will Not Work

When hitting sequence-length VRAM limits, engineers frequently attempt to scale existing distributed training strategies such as tensor parallelism, pipeline parallelism, or parameter sharding like ZeRO-3 vs FSDP. These configurations fail to resolve sequence-length memory pressure because none of them partition the sequence dimension itself.

Standard model-parallel and optimizer-sharding frameworks target distinct memory domains:

  • Tensor Parallelism (TP): Shards weight matrices and hidden dimensions across GPUs within a node, reducing parameter memory per GPU but retaining the full sequence length across attention and MLP operations.
  • Pipeline Parallelism (PP): Distributes transformer layers across sequential pipeline stages, which often worsens activation memory pressure because multiple microbatches must remain in flight to prevent pipeline bubbles.
  • ZeRO Stage 1 and Stage 2: Shard optimizer states and gradients across data-parallel ranks (Nd). The memory requirement per parameter is 4 + 12 / Nd bytes for Stage 1 and 2 + 14 / Nd bytes for Stage 2 in 16-bit mixed precision.
  • ZeRO Stage 3 / FSDP: Shards parameters down to 16 / Nd bytes per parameter during training.

While ZeRO and FSDP compress static states, activations sit on top of these states unsharded. Techniques like gradient checkpointing recompute activations during the backward pass to trade compute for memory, but at context lengths above 32K tokens, even the minimal stashed activation checkpoints exhaust per-GPU VRAM without sequence-level partitioning.

Parallelism TechniqueTarget Sharding DomainHandles Sequence ScalingVRAM Impact on Long Sequences
Tensor Parallelism (TP)Weights & Hidden Dimension (h)NoReduces per-GPU parameters but retains full sequence activations
Pipeline Parallelism (PP)Model Depth / Layers (L)NoIncreases peak activations due to multiple in-flight microbatches
ZeRO-1 / ZeRO-2 / FSDPOptimizer States & GradientsNoShrinks static state footprint; activation memory remains unsharded
Context Parallelism (CP)Sequence Dimension (s)YesDirectly divides per-GPU activation footprint across CP ranks

Context Parallelism vs. Sequence Parallelism

The terminology surrounding distributed sequence processing is often conflated across frameworks. In standard Megatron-LM implementations, sequence parallelism refers specifically to sharding activation tensors along the sequence dimension across the tensor-parallel group for operations outside the self-attention core (such as LayerNorm, Dropout, and MLP projections), which Korthikanti et al. report cuts activation memory by roughly 5x when combined with selective recomputation.

In contrast, context parallelism partitions the input sequence across an independent process group specifically for the attention mechanism itself. NVIDIA Megatron-Core defines context parallelism as designed to 'Split long sequences across GPUs for efficient long-context training', recommending its use for sequence lengths of 8K tokens and above.

When structuring a distributed run combining multiple parallelism dimensions, the total GPU allocation follows Megatron-Core's composition formula:

Total GPUs = TP * PP * CP * EP * DP

  • TP (Tensor Parallel): Shards hidden dimensions, constrained by intra-node NVLink bandwidth.
  • PP (Pipeline Parallel): Partitions model layers across nodes, trading communication latency for pipeline bubble overhead.
  • CP (Context Parallel): Partitions sequence length across ranks, directly dividing activation memory by the CP degree.
  • EP (Expert Parallel): Distributes Mixture-of-Experts routing across devices.
  • DP (Data Parallel): Replicates the model configuration across batches.

Ring Attention Demystified

Ring Attention provides an exact, distributed mechanism to compute full self-attention across multiple GPUs without ever gathering the global sequence into a single device's memory. Instead of performing an all-to-all broadcast of tokens, Ring Attention establishes a logical ring topology across all GPUs assigned to the context-parallel group.

The foundational principles of Ring Attention build upon the tiling strategy introduced in FlashAttention by Dao et al., submitted on 27 May 2022, which uses tiling to reduce the number of memory reads and writes between GPU high-bandwidth memory and on-chip SRAM while computing exact attention. By maintaining running softmax normalization statistics (the row maximum and running sum of exponentials), each GPU can compute partial attention outputs against local chunks and update them incrementally.

  1. Sequence Partitioning: The total sequence length s is split equally across N context-parallel GPUs, such that each GPU holds a local Query (Q), Key (K), and Value (V) block of length s / N.
  2. Local Block Computation: At step 0, each GPU computes local self-attention between its stationary Q block and its initial local K and V blocks.
  3. Asynchronous P2P KV Shift: Simultaneously, each GPU initiates an asynchronous non-blocking peer-to-peer send of its K and V blocks to the next rank in the ring (rank i -> rank (i+1) % N) and receives K and V blocks from the previous rank.
  4. Overlapped Kernel Execution: While the communication hardware transfers the next KV chunk over the interconnect, the GPU executes the attention kernel on the currently loaded chunk, updating local softmax accumulators.
  5. Ring Completion: After N - 1 communication hops, every rank has evaluated attention between its local Q block and the entire global sequence without ever materializing the global s-by-s attention matrix.

Ring Attention vs. DeepSpeed Ulysses (All-to-All)

When deploying context parallelism, engineering teams choose between two primary algorithmic families: Ring Attention and DeepSpeed Ulysses. Both approaches partition input tokens along the sequence dimension, but they execute attention redistribution through distinct communication patterns and collective operations.

DeepSpeed Ulysses, introduced by Jacobs et al. and submitted in September 2023, partitions input data along the sequence dimension and employs an efficient all-to-all collective communication for attention computation. Before computing attention, an all-to-all collective converts the sequence-partitioned tensor into a head-partitioned tensor, allowing each GPU to run standard attention on full sequence lengths for a subset of attention heads. After attention computation, a second all-to-all collective restores the sequence-partitioned layout for subsequent MLP layers.

Architectural FeatureDeepSpeed Ulysses (All-to-All)Ring Attention (P2P Ring)
Primary Communication PatternTwo All-to-All collectives per layerN-1 Point-to-Point (P2P) ring transfers per layer
Scaling ConstraintStrictly capped by number of attention heads (H >= CP)Scales arbitrarily with sequence length; no head cap
Grouped-Query Attention (GQA)Constrained by small KV head count (e.g. 8 KV heads)Unconstrained; supports arbitrary GQA configurations
Interconnect DependencyRequires high bisection bandwidth (NVLink or full-crossbar InfiniBand)Requires dedicated bi-directional ring bandwidth between adjacent ranks
Memory OverheadZero intermediate KV communication buffersRequires double-buffering for overlapped P2P KV transfers

The structural limitation of DeepSpeed Ulysses appears when training modern architectures that use Grouped-Query Attention (GQA). For example, a model with 8 KV heads cannot scale Ulysses beyond a CP degree of 8 without complex head-replication schemes. Conversely, Ring Attention scales independently of head counts, but naive ring implementations risk becoming communication-bound if peer-to-peer transfer latency exceeds local kernel execution time.

Scaling Constraints and Hardware Dependencies

Context parallelism is fundamentally an interconnect-bound optimization. As sequence lengths scale into hundreds of thousands of tokens, maintaining high Model Flops Utilization (MFU) requires matching the context-parallel degree to the physical fabric bandwidth.

Empirical scaling data published in the Meta Llama 3 architecture report illustrates how context parallelism scales across massive infrastructure clusters. Across the pre-training configurations in its Table 4, reported BF16 Model FLOPs Utilization is 43% on the smaller 8,192-GPU configuration at context-parallel degree 1, and on the larger 16,384-GPU configuration it falls from 41% at context-parallel degree 1 to 38% once context parallelism is raised to degree 16 for long sequences:

  • 8,192 NVIDIA H100 GPUs at CP degree 1 with 8,192-token sequences delivered 430 TFLOPs/GPU, corresponding to 43% BF16 MFU.
  • 16,384 NVIDIA H100 GPUs at CP degree 1, still on 8,192-token sequences, delivered 400 TFLOPs/GPU and 41% BF16 MFU.
  • 16,384 NVIDIA H100 GPUs scaled to CP degree 16 for the long-context stage maintained 380 TFLOPs/GPU and 38% BF16 MFU.

Hardware topology defines where context parallelism can be efficiently deployed. On single-node environments, Lyceum provides On-demand GPU VM configurations with 1 to 8 GPUs per instance. For SXM-based nodes (such as NVIDIA H100 or NVIDIA B200), NVLink provides up to 900 GB/s to 1.8 TB/s of bidirectional interconnect, easily saturating P2P and all-to-all exchanges. However, PCIe-only hardware like the NVIDIA L40S lacks NVLink support, making multi-GPU context parallelism on L40S nodes bottlenecked by host PCIe bus bandwidth.

When calculating per-device memory budgets for next-generation hardware like the NVIDIA B200, system architects must use the shipping HGX/DGX B200 specification of 180 GB usable HBM3e per GPU, since NVIDIA's DGX B200 documentation publishes 1,440 GB of total GPU memory across the eight GPUs in the system, rather than the 192 GB architectural chip capacity.

Large-Scale GPU Cluster for Long-Context Runs

Scaling context parallelism across sequences exceeding 128K tokens requires moving beyond single-node boundaries into dedicated multi-node clusters equipped with ultra-low-latency networking fabric. Without dedicated internode throughput, inter-device KV transfers stall CUDA execution pipelines, collapsing hardware efficiency.

For distributed training runs that require scalable context parallelism, Lyceum Technology offers Large-Scale GPU Cluster infrastructure designed specifically to eliminate communication bottlenecks:

  • Scale and Fabric: Multi-node allocations interconnected with non-blocking 400 Gb/s InfiniBand NDR networking to ensure full overlap of ring communication and compute kernels.
  • Rapid Allocation: Automated cluster provisioning within 28 seconds, minimizing setup overhead.
  • Flexible Workload Orchestration: Managed environments delivered on Slurm or Kubernetes, configured directly to team specifications.
  • Data Sovereignty: Hosted across European data centres in Paris and Finland under full EU regulatory and GDPR compliance.
  • Commercial Terms: Available on 3, 6, 12, or 24-month commitments or custom terms, backed by quote-based delivery with a 24-hour turnaround.

By aligning context-parallel algorithmic choices with non-blocking InfiniBand hardware, engineering teams can scale sequence lengths to 128K, 512K, and beyond while maintaining peak GPU throughput and eliminating sequence-driven OOM crashes.