The Dual Pathology: Routing and Representation Collapse

Sparse Mixture-of-Experts (MoE) architectures exist to break the coupling between model capacity and inference compute. By replacing dense feed-forward networks with a pool of parallel sub-networks and a lightweight gating router, an MoE scales its total parameter footprint while keeping active parameter FLOPs roughly constant per token. Published MoE models carry that split in their names: in Qwen3-235B-A22B, the figure after the A is the number of parameters that actually fire per token, a small fraction of the total the checkpoint holds. That is what a working router buys architecturally. However, realizing these theoretical gains during pre-training requires overcoming two distinct failure modes.

When training an MoE model from scratch or executing a continue-pretraining run, the router can destabilize into catastrophic failure patterns that waste hardware and degrade representation quality. Engineers frequently bundle these failures under the catch-all term collapse, but they represent two mechanically separate pathologies that require different telemetry to diagnose:

  • Routing collapse (Token starvation): The router converges onto a fixed, small subset of experts across all input tokens. Because unselected experts receive zero routing probability, they receive zero gradient updates during the backward pass. Over successive training steps, these starved experts freeze in their initial state while the active experts handle all workload. The training run continues to allocate full VRAM and perform all-to-all communication steps across GPUs, but the effective capacity drops to that of a small dense network.
  • Representation collapse (Functional homogenization): The router distributes tokens relatively evenly across the entire pool, but the experts fail to specialize. Instead of partitioning the feature space (e.g., syntax, mathematics, multilingual tokens), all experts learn near-identical parameter representations. The model exhibits uniform compute utilization, but architectural capacity collapses into redundant, duplicate sub-networks.

Both pathologies completely destroy the computational efficiency that justified the MoE architecture. For teams managing VRAM requirements and communication overhead, routing collapse is particularly punishing: you pay the full memory and network penalty for parameters that contribute nothing to the final loss. Preventing both requires active intervention across four distinct control levers: the auxiliary load-balancing loss, the router z-loss, the expert capacity factor, and the scope over which balance is computed.

Auxiliary Load-Balancing Loss: Tuning the Trade-Off

The standard mathematical defence against routing collapse is the auxiliary load-balancing loss (LBL), introduced by GShard and simplified by Switch Transformers. Rather than acting as a hard routing constraint, the auxiliary loss functions as a soft regularizer added directly to the primary language modelling objective. For an MoE layer with N experts and a batch of T tokens, the loss penalizes divergence between the actual routing assignments and uniform distribution.

The standard auxiliary balancing loss is formulated as the scaled dot product between the fraction of tokens dispatched to each expert and the mean gating probability assigned to that expert:

  1. Dispatch fraction (f_i): The proportion of tokens in the batch assigned to expert i, defined as f_i = (N / (K * T)) * sum(1(token_t selects expert_i)), where K is the number of active experts per token.
  2. Routing probability (P_i): The mean softmax probability assigned to expert i across all tokens in the batch, defined as P_i = (1 / T) * sum(s_{i,t}), where s_{i,t} is the router gate score.
  3. Auxiliary objective: L_balance = alpha * N * sum(f_i * P_i) for i in 1..N, where alpha is the balancing loss hyperparameter.

Because the dispatch fraction f_i is non-differentiable (derived from a discrete top-k argmax operation), gradients cannot flow through f_i directly. Instead, backpropagation acts entirely through the smooth routing probabilities P_i. Minimizing this product pushes both the discrete token assignments and the continuous gating distribution toward a uniform 1/N allocation.

The behaviour at either extreme of the coefficient is easier to state in prose than in a table, because the published evidence is qualitative on one side and a single swept value on the other. With no balancing term at all (alpha = 0), nothing opposes routing collapse, which is precisely the failure mode the auxiliary-loss and loss-free balancing literature exists to control. Switch Transformers used alpha = 10^-2 throughout their experiments, selected after sweeping alpha from 10^-1 to 10^-5 as a value large enough to ensure load balancing yet small enough not to overwhelm the primary cross-entropy objective. Smaller coefficients interfere less with the task gradient and leave more room for domain-driven specialization, but bite later and less firmly; larger ones hold load flat while measurably degrading model quality, and no source publishes a coefficient that is correct across architectures.

The auxiliary coefficient alpha is not a configuration toggle with a universally correct setting; it represents a direct trade-off between hardware utilization and model expressivity. If alpha is tuned too high, the balancing gradient overpowers the task gradient, forcing the model to route tokens uniformly even when specific tokens clearly belong to specialized experts. To diagnose this interaction, you must log L_balance as an independent telemetry curve alongside the cross-entropy task loss. If L_balance stays pinned at its theoretical minimum while task loss plateaus, your auxiliary loss is actively degrading model convergence.

Router Z-Loss: Numerical Stability vs. Distribution

A common operational mistake during MoE pre-training is attributing all loss spikes to load imbalance. In practice, training instability frequently originates from numerical divergence in the router itself. Because router logits enter an exponentiated softmax function, large logit magnitudes produce extreme values that cause catastrophic underflow or overflow when evaluated in 16-bit precision formats (fp16 or bf16).

To prevent logit drift without altering routing balance, ST-MoE introduced the router z-loss. The z-loss applies a direct penalty to the squared log-sum-exp of the unnormalized router logits (x):

  • Z-Loss Formulation: L_z = (c_z / B) * sum(log(sum(exp(x_{b,j})))^2), where B is the number of tokens, j iterates over all experts, and c_z weights the term inside the total loss. ST-MoE selected a small c_z on the basis of best post-pre-training model quality across a hyperparameter sweep.
  • Softmax Stabilization: By penalizing large values of log(sum(exp(x))), the router is prevented from pushing logits into high-magnitude regimes where 16-bit rounding turns differences into zeros or infinities. In ST-MoE's runs the z-loss stabilized the model without the quality damage caused by tighter constraints on activations and gradients.
  • Independence from Load: Z-loss exerts no direct pressure on expert distribution. A model can route perfectly uniformly and still suffer logit explosion, or hold stable logits while collapsing onto a single expert.

Maintaining stability across long pre-training runs requires treating router z-loss and load-balancing loss as distinct health metrics. When reviewing training telemetry, monitor three orthogonal signals: task cross-entropy, auxiliary balance loss, and z-loss magnitude. If your training run experiences sudden gradient norm spikes or non-finite values (NaNs) immediately after checkpoint resumption, the culprit is almost invariably router logit drift addressable via z-loss, not expert load imbalance.

Expert Capacity Factor and Token Dropping

While the auxiliary loss encourages routing balance, distributed execution frameworks historically required static tensor allocations to maintain synchronous GPU execution. To allocate fixed memory buffers for each expert on every device, frameworks implement an expert capacity factor (ECF). The capacity factor defines the maximum number of tokens an expert can accept within a single forward pass.

Expert capacity is calculated as: Capacity = ceil(ECF * (Total Tokens / Total Experts) * K). If the router assigns more tokens to an expert than its buffer allows, the excess tokens cannot be processed by that expert. Depending on the framework configuration, these overflow tokens are either dropped (bypassing the expert layer via the residual connection) or padded with dummy compute.

  1. ECF = 1.0 (Strict Capacity): Buffer size perfectly matches an idealized uniform distribution. A capacity factor greater than 1.0 exists precisely to create buffer for tokens that are not perfectly balanced across experts, so at 1.0 even minor imbalance produces dropped tokens that skip the expert and pass through the residual connection.
  2. Modest headroom above 1.0: Absorbs natural routing variance across batches. ST-MoE studies the capacity factor directly and shows that a train capacity factor above 1.0 sharply reduces the share of tokens dropped compared with a sub-1.0 setting, and the extra buffer is paid for in activation memory whether or not it is used.
  3. ECF >= 2.0 (Overprovisioned): Eliminates dropped tokens entirely during unbalanced phases, but inflates memory and compute overhead. Studies on ST-MoE-32B corroborate that high capacity factors do not improve fine-tuning quality.
  4. Dropless MoE (Dynamic Grouped Kernels): Some implementations discard static buffers entirely in favour of variable-length, block-sparse matrix multiplications, the dMoE strategy used in MegaBlocks, which the load-balancing-loss literature lists among the frameworks that compute balance at micro-batch scope. Dropless routing processes all tokens regardless of distribution, shifting the penalty of imbalance from dropped data to GPU execution stragglers.

If your architecture relies on capacity buffers, token drop rate is a mandatory per-layer metric. Published runs report drop rates falling to a small fraction of tokens once the auxiliary loss carries a high enough coefficient, but a low aggregate figure can still conceal a severe local failure where one deep layer is discarding a large share of its incoming tokens. For detailed memory allocation and tensor sizing strategies when configuring training runs, review our guide on preventing OOM errors in deep learning infrastructure.

Balance Measurement Scope: Global-Batch vs Micro-Batch

The most critical yet frequently overlooked parameter in MoE training is the scope over which expert load is measured. In standard distributed data-parallel (DDP) or fully sharded data parallel setups, training frameworks calculate the auxiliary balancing loss independently within each GPU's local micro-batch, subsequently averaging the scalar loss across data-parallel ranks.

When training billion-scale language models, a single micro-batch on a GPU may contain only a few thousand tokens (often just 1 to 4 packed sequences). Calculating the auxiliary loss at the micro-batch level creates an overly strict constraint: it forces the router to distribute tokens evenly within each individual sequence. If a micro-batch contains a domain-specific sequence (such as Python source code or LaTeX mathematics), micro-batch balancing actively punishes the router for sending all code tokens to a dedicated programming expert.

  • Micro-Batch Balancing (LBL_micro): Calculates dispatch fraction f_i and gating probability P_i solely within each parallel group, then averages the scalar across groups. Because a micro-batch for a billion-scale run holds very few sequences, this is almost a sequence-level constraint and inhibits expert specialization.
  • Global-Batch Balancing (LBL_global): Synchronizes expert selection counts across all data-parallel ranks before computing the auxiliary loss. Balances routing across the broader corpus rather than inside single sequences.
  • Communication Cost: Global-batch balancing adds one extra communication step that synchronizes the per-expert selection frequency f_i across micro-batches before the loss is computed, which is why the authors treat it as a cheap addition to the training step.
  • Empirical Validation: Experiments on MoE LLMs of up to 42.8 billion total parameters trained on 400 billion tokens found global-batch LBL yields excellent performance gains in both pre-training perplexity and downstream tasks, and greatly improves domain specialization.

Global-batch balancing resolves the fundamental conflict between expert specialization and hardware efficiency. By synchronizing expert assignment frequencies across ranks or accumulating them across gradient accumulation steps via an internal count buffer, the model achieves macro-level hardware balance without penalizing micro-level sequence specialization. For multi-node cluster synchronization mechanics, see our distributed training guide.

Auxiliary-Loss-Free Balancing: The Bias Alternative

While global-batch auxiliary loss mitigates the specialization penalty, traditional auxiliary losses still inject non-task gradients directly into the routing network. These interference gradients can pull the router parameters away from optimal token-to-expert representations. To eliminate gradient conflict entirely, researchers developed auxiliary-loss-free balancing strategies.

Auxiliary-loss-free balancing removes the regularization penalty from the loss function completely. Instead, it introduces a dynamic, per-expert bias term (b_i) added directly to the routing scores during top-k selection:

  1. Biased Routing Decision: For input token u_t and expert centroid e_i, the unnormalized affinity score is s_{i,t} = G(u_t^T * e_i). Top-k expert selection is evaluated over (s_{i,t} + b_i), routing the token to experts with the highest biased scores.
  2. Unbiased Weighting: The bias b_i is used solely for the discrete dispatch decision. When computing the weighted combination of expert outputs, the original gating weight s_{i,t} (or its normalized softmax/sigmoid value) is used without the bias term.
  3. Dynamic Feedback Loop: At the end of each training step, the framework monitors the total token allocation c_i per expert across the global batch. If an expert receives more than the average token load (c_i > c_mean), its bias is decremented by a small step (b_i = b_i - u * sign(c_i - c_mean)). If an expert is starved, its bias is incremented.
Balancing StrategyMeasured Load Violation (MaxVio_global)Reported Validation PerplexityImplementation Mechanism
Auxiliary loss, alpha = 0.001 (1B model)0.729.56Loss term added to the backward graph
Auxiliary loss, alpha = 0.001 (3B model)0.527.97Loss term added to the backward graph, optionally with a cross-rank count all-reduce for global-batch scope
Loss-free bias (3B model)0.047.92Per-step scalar bias update on the gating scores used for top-K selection only

Evaluated on MoE models of up to 3 billion parameters trained on up to 200 billion tokens, auxiliary-loss-free balancing achieved both better perplexity and much better global load balance than the auxiliary-loss-controlled baseline. Because the bias adjustments are applied to the top-K selection according to recent load rather than through the loss function, the method produces no interference gradients, so the parameter updates come from the language modelling objective alone. Teams should view loss-free bias routing and global-batch auxiliary loss as two viable architectural options depending on whether their training framework supports decoupled routing biases.

Expert Parallelism as a Distributed Systems Problem

Ultimately, load balancing in an MoE model is not merely a theoretical optimization problem; it is a distributed systems throughput bottleneck. When scaling MoE pre-training across multiple nodes, expert parallelism (EP) shards different experts across different physical GPUs. In an EP configuration, every single MoE layer requires an all-to-all collective communication operation: GPUs must exchange assigned tokens before expert forward execution, and exchange intermediate activations after expert computation distributed runs.

Because the all-to-all operation is synchronous, the execution time of the entire distributed layer is dictated by the most heavily loaded GPU shard. Every other rank sits at the communication barrier until the overloaded expert shard finishes, so step time scales with the max-to-mean load ratio across experts rather than with the average load. This is the efficiency argument the MoE literature makes for balancing at all: when experts are spread across devices, imbalanced expert utilization heavily slows the forward pass, and on large clusters that shows up as degraded model FLOPs utilization long before it shows up in validation loss.

  • Single-Node Debugging: Before launching a multi-node pre-training run, isolate router dynamics, z-loss stability, and capacity factors on a single node. An on-demand GPU VM with 1 to 8 GPUs is enough, and a high-bandwidth NVLink domain (600 GB/s bidirectional per GPU on NVIDIA A100, 900 GB/s on H100 and H200, up to 1.8 TB/s on B200) lets you profile gating distributions without an inter-node network in the path. Note that NVIDIA L40S instances have no NVLink at all, and that an NVLink domain covers one baseboard of 4 or 8 GPUs; anything wider is InfiniBand.
  • Managed Execution: For containerized pre-training jobs, a serverless training product that provisions and starts a submitted run in under 60 seconds, pulling images from Docker Hub, ECR, or GAR with S3-compatible object storage, removes the cluster-setup step from the debug loop.
  • Large-Scale Cluster Scaling: Multi-node expert parallelism is bounded by inter-node bandwidth, so the fabric is the specification that matters. Lyceum's Large-Scale GPU Cluster runs on 400 Gb/s InfiniBand NDR, the per-port rate NVIDIA publishes for its NDR platform, which is the rate of one link rather than a per-node aggregate. It is managed on Slurm or Kubernetes with 3, 6, 12, or 24-month terms, deployed across European data centres in Paris and Finland.

By aligning your architectural regularizers (global-batch balancing and router z-loss) with high-bandwidth interconnect infrastructure, you ensure that every expert in the network receives sufficient gradient updates while maximizing cluster throughput across the entire pre-training lifecycle.