Plan on the differences, not on the generation
Most cluster plans start from the wrong question: what does the new generation cost per GPU-hour. That number is the least useful figure in the decision. A training team does not buy GPU-hours, it buys wall-clock time to a finished run, and the parts that determine that time are not the FLOPS count. They are how much memory sits on each device, how the devices connect, what the rack draws in power and heat, and what unit the vendor actually sells. Plan a B300 or GB200 cluster as a faster B200 cluster and you will size the wrong thing.
Check CUDA, driver, kernel, communication library and container compatibility when moving hardware generations. Existing software may need updated builds or tuning. Hardware topology and software support both affect the decision.
- Memory per device: changes how the model shards and how many nodes the job needs.
- Interconnect topology: decides whether extra memory per device actually reduces the node count.
- Power and cooling envelope: decides which facilities can host the part at all.
- Purchasing unit: node-level versus rack-scale, which decides what a six-to-twelve-node buyer can actually order.
The node-level reference points are NVIDIA's DGX B200 and DGX B300 systems, both eight-GPU nodes in 10 rack units. They are the cleanest apples-to-apples pair for sizing a node-quantity cluster.
What more memory per device changes
NVIDIA’s DGX B200 user guide specifies 1,440 GB across 8 GPUs. The DGX B300 user guide specifies 8 × 288 GB, about 2.3 TB. These are complete-system specifications. Do not combine a generic 192 GB B200 listing with DGX B200’s total. NVIDIA marketing pages also show different rounded B300 totals, so confirm the quoted configuration and usable memory.
| System specification | DGX B200 user guide | DGX B300 user guide |
|---|---|---|
| GPU count | 8 | 8 |
| Total GPU memory | 1,440 GB | 8 × 288 GB, about 2.3 TB |
| Cluster interfaces | 8 ConnectX-7 cards, up to 400 Gb/s each | 8 ConnectX-8 cards, up to 800 Gb/s each |
| Configuration check | Confirm quoted memory and power envelope | Guide lists 14.5 kW consumption and 15 kW system maximum in its power section |
More memory can reduce the number of devices needed for a given workload, but it does not guarantee fewer failures or a shorter run. Check the model’s memory use, parallelism and restart behaviour on the proposed topology.
Interconnect and how jobs shard across nodes
Extra memory per device only cuts the node count if the interconnect keeps up. Two fabrics matter, and they solve different problems. Scale-up is the GPU-to-GPU fabric inside a node or rack: fifth-generation NVLink delivers 1,800 GB/s per GPU, and the GB200 NVL72 connects 72 Blackwell GPUs into a single NVLink domain moving 130 TB/s. Scale-out is the fabric between nodes, where InfiniBand carries the collectives that do not fit inside the NVLink domain.
| Fabric | Domain | Bandwidth |
|---|---|---|
| NVLink 5 (scale-up, per GPU) | Within a node or NVLink rack | 1,800 GB/s per GPU |
| NVLink Switch (GB200 NVL72) | 72 GPUs in one liquid-cooled rack | 130 TB/s aggregate |
| ConnectX-8 (scale-out, per GPU) | Between nodes | Up to 800 Gb/s InfiniBand |
| Lyceum cluster fabric | Between nodes, managed | 400 Gb/s InfiniBand NDR |
The practical consequence for shard planning: tensor parallelism wants to live inside the NVLink domain, data and pipeline parallelism tolerate the InfiniBand hop. A rack-scale NVL72 system gives you a 72-GPU tensor-parallel domain in one enclosure; a node-level B300 cluster gives you eight-GPU domains joined by InfiniBand. Both train large models. They shard differently, and your parallelism strategy should match the topology you actually buy.
Rack-scale GB200 against node-level B300
Lyceum lists GB200 and GB300 as quote-based cluster configurations. Ask which complete system and allocation are offered. A hardware family listing does not establish a full-rack minimum order or current capacity at a particular European site.
- B300: confirm the complete node configuration and the allocation the provider sells
- GB300 NVL72: a 72-GPU liquid-cooled system with 130 TB/s aggregate NVLink and about 20 TB GPU memory
- GB200 NVL72: a liquid-cooled design combining 36 Grace CPUs and 72 Blackwell GPUs
Compare the allocation the supplier will actually sell: GPU count, NVLink domain, network and term. A rack-scale design describes the hardware, while the commercial allocation depends on the service and quote.
Where B300 and GB200 can be obtained in Europe
Lyceum lists B300 across VM, dedicated inference, managed training and cluster offerings. GB200 and GB300 require a cluster quote. Confirm the mode, quantity, site and start date. The on-demand dashboard showed no B300 capacity on 1 October 2026; a catalogue listing is not proof of stock.
- Available capacity: check the live offering for the exact GPU, site and mode
- New machines: the standard planning lead time is around 4 weeks, subject to a confirmed quote
- Reservations: the stated minimum is one month on one server; this does not apply to on-demand VMs
- Clusters: obtain a written configuration, capacity allocation and delivery date
Power and cooling decide which facilities qualify
Use the exact system’s electrical and cooling specification for facility planning. The DGX B300 guide lists 14.5 kW consumption and 15 kW maximum in its power section. A 12-node allocation at 14.5 kW represents 174 kW of system load before networking, storage and cooling overhead. This is sizing arithmetic, not measured workload consumption.
| System | Cooling and sizing check |
|---|---|
| DGX B200 | Confirm the quoted node’s airflow, electrical maximum and rack layout |
| DGX B300 | Airflow and AC or DC supply requirements; confirm the guide and quoted configuration |
| GB200 NVL72 | Liquid-cooled rack, facility integration and rack power requirements |
| GB300 NVL72 | Liquid-cooled rack, facility integration and rack power requirements |
Choosing available now over faster later
Use a dated benchmark result with its hardware count, software, precision and target quality, then validate the workload you plan to run. MLPerf Training provides a structured comparison, but a benchmark configuration does not establish a supplier’s available capacity.
The cost model follows from that. An extra month waiting for the newer part is a month of runway spent, and a run that finishes six weeks earlier on the previous generation is six weeks of iteration your competitors do not get. Our cost per training run calculator prices the run this way, on total time to result rather than hourly rate. On Lyceum, Large-Scale GPU Cluster runs on terms of 3, 6, 12 or 24 months or custom, quote-based with a 24-hour response, and capacity can be added or removed with 2 to 3 weeks notice with no penalty for scaling down.
- Which site: which facility hosts the nodes, and does it have the power and cooling for the part?
- What lead time: firm date for the machines to be running, not a catalogue listing.
- What minimum commitment: term, quantity and notice period for scaling down.