Plan on the differences, not on the generation

Most cluster plans start from the wrong question: what does the new generation cost per GPU-hour. That number is the least useful figure in the decision. A training team does not buy GPU-hours, it buys wall-clock time to a finished run, and the parts that determine that time are not the FLOPS count. They are how much memory sits on each device, how the devices connect, what the rack draws in power and heat, and what unit the vendor actually sells. Plan a B300 or GB200 cluster as a faster B200 cluster and you will size the wrong thing.

Check CUDA, driver, kernel, communication library and container compatibility when moving hardware generations. Existing software may need updated builds or tuning. Hardware topology and software support both affect the decision.

  • Memory per device: changes how the model shards and how many nodes the job needs.
  • Interconnect topology: decides whether extra memory per device actually reduces the node count.
  • Power and cooling envelope: decides which facilities can host the part at all.
  • Purchasing unit: node-level versus rack-scale, which decides what a six-to-twelve-node buyer can actually order.

The node-level reference points are NVIDIA's DGX B200 and DGX B300 systems, both eight-GPU nodes in 10 rack units. They are the cleanest apples-to-apples pair for sizing a node-quantity cluster.

What more memory per device changes

NVIDIA’s DGX B200 user guide specifies 1,440 GB across 8 GPUs. The DGX B300 user guide specifies 8 × 288 GB, about 2.3 TB. These are complete-system specifications. Do not combine a generic 192 GB B200 listing with DGX B200’s total. NVIDIA marketing pages also show different rounded B300 totals, so confirm the quoted configuration and usable memory.

System specificationDGX B200 user guideDGX B300 user guide
GPU count88
Total GPU memory1,440 GB8 × 288 GB, about 2.3 TB
Cluster interfaces8 ConnectX-7 cards, up to 400 Gb/s each8 ConnectX-8 cards, up to 800 Gb/s each
Configuration checkConfirm quoted memory and power envelopeGuide lists 14.5 kW consumption and 15 kW system maximum in its power section

More memory can reduce the number of devices needed for a given workload, but it does not guarantee fewer failures or a shorter run. Check the model’s memory use, parallelism and restart behaviour on the proposed topology.

Interconnect and how jobs shard across nodes

Extra memory per device only cuts the node count if the interconnect keeps up. Two fabrics matter, and they solve different problems. Scale-up is the GPU-to-GPU fabric inside a node or rack: fifth-generation NVLink delivers 1,800 GB/s per GPU, and the GB200 NVL72 connects 72 Blackwell GPUs into a single NVLink domain moving 130 TB/s. Scale-out is the fabric between nodes, where InfiniBand carries the collectives that do not fit inside the NVLink domain.

FabricDomainBandwidth
NVLink 5 (scale-up, per GPU)Within a node or NVLink rack1,800 GB/s per GPU
NVLink Switch (GB200 NVL72)72 GPUs in one liquid-cooled rack130 TB/s aggregate
ConnectX-8 (scale-out, per GPU)Between nodesUp to 800 Gb/s InfiniBand
Lyceum cluster fabricBetween nodes, managed400 Gb/s InfiniBand NDR

The practical consequence for shard planning: tensor parallelism wants to live inside the NVLink domain, data and pipeline parallelism tolerate the InfiniBand hop. A rack-scale NVL72 system gives you a 72-GPU tensor-parallel domain in one enclosure; a node-level B300 cluster gives you eight-GPU domains joined by InfiniBand. Both train large models. They shard differently, and your parallelism strategy should match the topology you actually buy.

Rack-scale GB200 against node-level B300

Lyceum lists GB200 and GB300 as quote-based cluster configurations. Ask which complete system and allocation are offered. A hardware family listing does not establish a full-rack minimum order or current capacity at a particular European site.

  • B300: confirm the complete node configuration and the allocation the provider sells
  • GB300 NVL72: a 72-GPU liquid-cooled system with 130 TB/s aggregate NVLink and about 20 TB GPU memory
  • GB200 NVL72: a liquid-cooled design combining 36 Grace CPUs and 72 Blackwell GPUs

Compare the allocation the supplier will actually sell: GPU count, NVLink domain, network and term. A rack-scale design describes the hardware, while the commercial allocation depends on the service and quote.

Where B300 and GB200 can be obtained in Europe

Lyceum lists B300 across VM, dedicated inference, managed training and cluster offerings. GB200 and GB300 require a cluster quote. Confirm the mode, quantity, site and start date. The on-demand dashboard showed no B300 capacity on 1 October 2026; a catalogue listing is not proof of stock.

  • Available capacity: check the live offering for the exact GPU, site and mode
  • New machines: the standard planning lead time is around 4 weeks, subject to a confirmed quote
  • Reservations: the stated minimum is one month on one server; this does not apply to on-demand VMs
  • Clusters: obtain a written configuration, capacity allocation and delivery date

Power and cooling decide which facilities qualify

Use the exact system’s electrical and cooling specification for facility planning. The DGX B300 guide lists 14.5 kW consumption and 15 kW maximum in its power section. A 12-node allocation at 14.5 kW represents 174 kW of system load before networking, storage and cooling overhead. This is sizing arithmetic, not measured workload consumption.

SystemCooling and sizing check
DGX B200Confirm the quoted node’s airflow, electrical maximum and rack layout
DGX B300Airflow and AC or DC supply requirements; confirm the guide and quoted configuration
GB200 NVL72Liquid-cooled rack, facility integration and rack power requirements
GB300 NVL72Liquid-cooled rack, facility integration and rack power requirements

Choosing available now over faster later

Use a dated benchmark result with its hardware count, software, precision and target quality, then validate the workload you plan to run. MLPerf Training provides a structured comparison, but a benchmark configuration does not establish a supplier’s available capacity.

The cost model follows from that. An extra month waiting for the newer part is a month of runway spent, and a run that finishes six weeks earlier on the previous generation is six weeks of iteration your competitors do not get. Our cost per training run calculator prices the run this way, on total time to result rather than hourly rate. On Lyceum, Large-Scale GPU Cluster runs on terms of 3, 6, 12 or 24 months or custom, quote-based with a 24-hour response, and capacity can be added or removed with 2 to 3 weeks notice with no penalty for scaling down.

  1. Which site: which facility hosts the nodes, and does it have the power and cooling for the part?
  2. What lead time: firm date for the machines to be running, not a catalogue listing.
  3. What minimum commitment: term, quantity and notice period for scaling down.