AI This article was created with the help of AI.

The 4-Bit Hardware Leap: NVFP4 vs MXFP4

Before evaluating 4-bit training workflows, we must separate NVFP4 pretraining from post-training quantization. Compressing a converged checkpoint to 4-bit weights for serving (such as standard AWQ or GPTQ workflows) trades minimal perplexity for lower inference memory bandwidth. In contrast, 4-bit pretraining executes both the forward and backward GEMM operations directly inside hardware tensor cores at sub-byte precision throughout the entire run. This requires native silicon support that does not exist on Hopper or earlier architectures, rendering previous-generation techniques like FP8 training on NVIDIA H100 hardware the baseline rather than the ceiling.

The core numerical challenge in sub-byte training is preserving dynamic range and gradient fidelity. NVFP4 addresses this through an E2M1 format (1 sign bit, 2 exponent bits, and 1 mantissa bit), representing discrete values between -6.0 and +6.0. Rather than applying a single coarse scaling factor across an entire tensor, NVFP4 organizes elements into compact micro-blocks of 16 contiguous values. Each 16-element micro-block shares an 8-bit FP8 E4M3 scale factor, which is paired with a secondary, global FP32 per-tensor scalar. This two-level scaling mechanism enables the format to adapt dynamically to localized activation spikes.

In contrast, the Open Compute Project (OCP) MXFP4 specification uses a block size of 32 elements governed by an unsigned 8-bit power-of-two scale factor (E8M0). While power-of-two scaling simplifies hardware shift logic, it forces scale values to snap to discrete powers of two, widening quantization error across wide dynamic distributions. NVIDIA reports an average quantization mean squared error (MSE) of 0.08 for 16-element E4M3 scaling, and describes power-of-two E8M0 scaling as producing larger overall block errors because the scale snaps to the nearest power of two. That precision delta is the difference between stable loss progression and divergent gradients during large-scale pretraining.

FormatMicro-Block SizeBlock Scale RepresentationGlobal Tensor ScaleAverage Quantization MSE (NVIDIA)
NVFP416 elementsFP8 (E4M3) non-power-of-twoFP32 scalar0.08
OCP MXFP432 elementsE8M0 power-of-twoNone (single tier)Not published; NVIDIA reports larger block errors

On Blackwell silicon, these micro-block groupings and dynamic scaling operations execute directly inside fifth-generation Tensor Cores. The hardware handles scale alignment and sub-byte unpacking without falling back to software emulation, making NVFP4 a viable execution target for multi-node training clusters.

The NVFP4 Pretraining Recipe: What Stays in Higher Precision

Running 4-bit training successfully across long token horizons is not simply a matter of casting every tensor to NVFP4. The published recipe from NVIDIA (arXiv:2509.25149, submitted 29 September 2025) was validated by pretraining a 12-billion-parameter hybrid Mamba-Transformer model on 10 trillion tokens, a run NVIDIA reports converged stably with downstream accuracy comparable to an FP8 baseline. The central takeaway from this work is that maintaining loss stability requires keeping a deliberate subset of operations and state variables in higher precision.

First, master weights, optimizer states (such as Adam's first and second moments), and gradient accumulations remain strictly in FP32 or BF16. Converting optimizer states to 4 bits destroys the running momentum statistics necessary to navigate non-convex loss surfaces. Second, tensor-parallel cross-entropy reductions, layer normalizations, and attention softmax calculations remain in higher precision to prevent catastrophic underflow and overflow in sensitive activation channels.

  • Master weights and Adam optimizer states: Maintained in FP32/BF16 to preserve gradient accumulation accuracy.
  • Two-level tensor scaling: 16-element E4M3 micro-block scaling combined with per-tensor FP32 normalization.
  • Randomized Hadamard transforms: Applied to activation matrices before quantization to rotate outlier channels into uniform distributions.
  • Stochastic rounding on backward GEMMs: Prevents systematic gradient underflow when small updates map to narrow 4-bit buckets.
  • Selective precision fallbacks: Embedding layers, normalization blocks, and final transformer blocks stay in BF16/FP8 to protect output logits.

Activation outliers represent another critical failure mode in narrow numerical formats. Extreme outlier features in deep transformer layers can dominate an entire micro-block, forcing non-outlier values to round to zero. The NVFP4 recipe addresses this by applying a Random Hadamard Transform, an orthogonal rotation applied to the wgrad GEMM operands before quantization, smoothing peak magnitudes without altering the model architecture. Stochastic rounding is then applied to gradients: each scaled value is rounded probabilistically to one of the two nearest representable NVFP4 values, with the probability inversely proportional to the distance, so the expected value of the quantized tensor matches the original and no systematic rounding bias suppresses small gradient updates.

The Real Throughput Math: Dense vs Sparse TFLOPS

A common misconception among infrastructure teams is that dropping from 8-bit to 4-bit arithmetic doubles training throughput automatically. In practice, several structural overheads prevent linear scaling. The two-level scaling scheme introduces memory overhead: an 8-bit E4M3 scale factor for every 16 4-bit elements adds 0.5 bits per parameter, while the global FP32 scalar amortizes across the tensor. This brings the effective bit-width of NVFP4 to roughly 4.5 to 4.6 bits per element, compared to 4.25 bits for MXFP4.

When evaluating raw compute capacity across Blackwell hardware, advertised peak numbers must be separated into dense execution and structured sparsity. NVIDIA data sheets publish 2:4 sparse compute figures by default, but production LLM pretraining almost exclusively uses dense matrix multiplication. On an HGX B200 configuration, peak NVFP4 compute reaches 9 PFLOPS dense and 18 PFLOPS sparse per GPU, while the liquid-cooled GB200 reaches 10 PFLOPS dense and 20 PFLOPS sparse per GPU. For the Blackwell Ultra generation, NVIDIA's GB300 NVL72 specification table lists rack-scale FP4 Tensor Core throughput both with and without sparsity, and the gap between the two is narrower than the doubling buyers usually assume.

GPU Form FactorArchitecture GenerationDense NVFP4 ComputeSparse NVFP4 Compute (2:4)Architectural Memory Capacity
NVIDIA B200 (HGX)Blackwell9 PFLOPS18 PFLOPS192 GB HBM3e (180 GB usable in HGX)
NVIDIA GB200Blackwell10 PFLOPS20 PFLOPS192 GB HBM3e (186 GB in NVL72)
NVIDIA B300 (HGX)Blackwell Ultra15 PFLOPS20 PFLOPS288 GB HBM3e at 8 TB/s
NVIDIA GB300 (NVL72)Blackwell Ultra15 PFLOPS20 PFLOPS288 GB HBM3e per GPU

Memory capacity also requires precise contextualization when sizing workloads. While 192 GB is NVIDIA's architectural HBM3e capacity for the B200 part, NVIDIA's own HGX reference architecture tables list 180 GB of HBM3e per B200 SXM GPU and 1.44 TB per eight-GPU node, and GB200 NVL72 configurations expose 186 GB per GPU. For the B300, NVIDIA specifies 288 GB of HBM3e per GPU at up to 8 TB/s memory bandwidth. Infrastructure leads planning memory allocations can explore detailed memory breakdowns in our B200 memory architecture analysis.

The FP8 Baseline: A Grounded Scaling Anchor

Because publicly validated NVFP4 pretraining benchmarks are currently limited to isolated laboratory runs, engineering teams should anchor their throughput expectations to mature FP8 scaling data rather than synthetic inference benchmarks. FP8 training execution captures the real-world overheads of mixed-precision regimes, including scale factor synchronization, kernel launch latency, and activation casting.

Empirical data from NVIDIA NeMo Framework benchmarks across NVIDIA H100 clusters demonstrates how speedup over BF16 scales with parameter count. NVIDIA reports that its current scaling FP8 recipe on H100 delivers speedups ranging from 1.30x on smaller dense models like Llama 3 8B up to 1.53x on Llama 3.1 405B, with the gain rising as computational density grows. On NVIDIA DGX B200 hardware, hardware-accelerated MXFP8 achieves consistent speedups between 1.28x and 1.37x across model sizes, reflecting the efficiency of block-level scaling on fifth-generation Tensor Cores.

  • Llama 3 8B on H100 (NeMo, current scaling FP8): 1.30x training speedup over BF16 baseline.
  • Llama 3 70B on H100 (NeMo, current scaling FP8): 1.43x training speedup over BF16 baseline.
  • Llama 3.1 405B on H100 (NeMo): 1.53x training speedup over BF16 baseline.
  • MXFP8 on DGX B200 (NeMo): 1.28x to 1.37x speedup across dense architectures.

These figures illustrate why NVFP4 will not deliver a speedup over BF16 that tracks the bit-width ratio during training. When calculating job completion times, teams should measure sustained tokens per second per GPU alongside time to target loss, accounting for the uncompressed communication overheads inherent to distributed clusters.

Running the Convergence Validation Protocol

Before committing to multi-month cluster reservations, training teams must validate NVFP4 numerical stability on their specific model architectures. Standard loss curves can appear healthy in early steps before collapsing when learning rate schedules decay. A rigorous pilot protocol eliminates this risk by benchmarking NVFP4 against a fixed higher-precision run under identical conditions.

  1. Establish identical seed, data order, and batch parameters: Run a baseline BF16 or FP8 run over a fixed short token budget with the same random seeds and identical data loader sharding.
  2. Instrument layer-by-layer quantization telemetry: Monitor the maximum absolute values (amax) and scale factor drift across micro-blocks to detect outlier accumulation early.
  3. Compare intermediate loss trajectories: Evaluate validation loss at a fixed mid-training checkpoint rather than waiting for final convergence. Ensure the loss delta remains within predefined limits.
  4. Execute downstream task evaluation checkpoints: Run zero-shot and few-shot evals on domain-specific benchmarks at regular intervals to catch latent quality degradation.
  5. Enforce a strict abort criterion: Agree a hard stopping rule before the run starts, expressed as a maximum sustained validation loss gap against the baseline over an agreed step window, so unviable runs terminate automatically.

If a loss gap persists late in training, NVIDIA's NVFP4 pretraining report observes that switching from FP4 to higher precision towards the end of training closes that gap, and recommends making the switch shortly before the onset of learning rate decay for full loss recovery, or at the very end for a smaller improvement with minimal effect on runtime. This hybrid approach captures the throughput gains of NVFP4 during the bulk of pretraining while ensuring the final checkpoint tracks higher-precision accuracy.

Lyceum's European Blackwell Footprint

Executing large-scale AI pretraining in Europe requires infrastructure that combines high-performance silicon with strict data sovereignty. At Lyceum, we provide enterprise access to NVIDIA B200, NVIDIA B300, NVIDIA GB200, and NVIDIA GB300 compute deployed exclusively in European data centres across Paris and Finland.

Every compute node in our footprint operates under European jurisdiction, supporting your GDPR obligations without exposure to foreign extra-territorial data access laws. We enforce a zero data retention architecture across our platform: prompts and outputs are processed but not stored.

  • European data residency: Facilities situated in Paris and Finland within the EU/EEA.
  • Sovereign compliance: Fully aligned with the GDPR.
  • Zero data retention: Customer datasets, model weights, and intermediate states remain strictly confidential.
  • Open-stack transparency: Built on open components like vLLM and NVIDIA Dynamo, avoiding proprietary black-box orchestration layers.

By coupling open-stack orchestration with dedicated Blackwell silicon, engineering teams maintain full compiler-level and kernel-level control over their distributed training workloads without operational friction.

Matching Workload to Supply: VMs, Serverless, and Clusters

Selecting the appropriate delivery model depends on the scale and duration of the training phase. For initial kernel profiling and convergence validation, multi-GPU distributed training pilots fit naturally onto our On-demand GPU VM offering. These instances provide raw SSH access to 1 to 8 GPUs with native NVLink interconnects, 18-second provisioning, per-second billing, and no data egress fees, allowing teams to test NVFP4 kernels without upfront commitments.

For repeatable fine-tuning and intermediate training jobs, Serverless Training provides managed container execution with job startup times under 60 seconds. Teams can submit training tasks with custom Docker images from ECR, GAR, or Docker Hub, backed by high-throughput S3-compatible object storage. For large-scale pretraining runs spanning 8 to 8,000 GPUs, our Large-Scale GPU Cluster delivers dedicated Slurm or Kubernetes environments backed by 400 Gb/s InfiniBand NDR non-blocking networking fabrics on 3, 6, 12, or 24-month terms.

Delivery ModeHardware OptionsInterconnect FabricProvisioning TimeTarget Workload
On-demand GPU VMNVIDIA B200, NVIDIA B300NVLink (intra-node)18 secondsKernel profiling, small-scale convergence pilots
Serverless TrainingNVIDIA B200, NVIDIA B300NVLink (node-level)< 60 secondsContainerized fine-tuning and automated training jobs
Large-Scale GPU ClusterB200, B300, GB200, GB300400 Gb/s InfiniBand NDR28 seconds (contracted)Multi-node pretraining (8 to 8,000 GPUs)

Network bandwidth is especially critical in 4-bit regimes. While NVFP4 accelerates matrix math, gradient all-reduce and optimizer state traffic do not shrink proportionally when using standard ZeRO-3 or FSDP sharding. Our InfiniBand NDR fabric provides the throughput necessary to keep Blackwell Tensor Cores fully saturated. Teams planning long-term compute roadmaps can consult our B200 deployment economics guide for structural cost analysis.

Before committing capital to extended training contracts, validate your convergence metrics on real silicon. Contact our infrastructure engineering team to request a customized Blackwell cluster quote and configure your same-seed validation benchmark.