Training
9 articlesCompute Budgeting After the Pre-Training Plateau
Frontier compute is shifting from pre-training toward post-training and inference-time reasoning. That changes how an enterprise budgets GPU capacity: fewer giant runs, more RL loops and evaluation, and a different…
Training Goodput at 10,000 GPUs: Failures, MFU and Honest Throughput
Llama 3 pre-training on 16,384 GPUs saw 466 interruptions in 54 days. At that failure rate, the number that matters is goodput, not peak FLOPS. This explains MFU, failure modes,…
800 VDC Power: Preparing Training Halls for Megawatt AI Racks
Rack power went from about 40 kW in the Hopper era to roughly 120 kW with Blackwell, and 800 VDC distribution arrives with megawatt racks from 2027. This explains what…
HBM4 and the Memory Wall: What It Means for 2027 Training Clusters
HBM4 doubles the memory interface to 2048 bits and entered mass production in early 2026. Because large-model training is bandwidth-bound more often than FLOPS-bound, that step decides achieved utilisation, cluster…
Vera Rubin NVL144: What the 2026 Training Platform Changes for Cluster Design
Vera Rubin NVL144 keeps the rack as the scale-up domain but raises memory, interconnect and power together. HBM4, NVLink 6 and ConnectX-9 change how many racks a training run needs,…
FP8 to FP4: How Low-Precision Training Reshapes Cluster Sizing
FP8 pretraining is the 2026 default and NVFP4 4-bit recipes are validated to 120B scale with FP8-matching accuracy, doubling arithmetic and halving memory on Blackwell-class silicon. Size clusters in tokens-per-day…
Blackwell Ultra to Vera Rubin to Feynman: The 2026–2028 AI Training Cluster Roadmap
NVIDIA now ships one AI architecture per year: Blackwell Ultra today, Vera Rubin from H2 2026, Rubin Ultra in 2027, Feynman in 2028. For most training clusters the deciding factor…
Reference Architecture: Sovereign AI Cluster (Scalable Unit)
This reference architecture specifies an in-country, DPDP-aware GPU cluster built from a repeatable Scalable Unit (SU): a group of 8× H200 nodes joined by NDR/XDR InfiniBand, with shared parallel storage…
Reference Architecture: 8× H200 On-Prem AI Training Node
This reference architecture specifies a single 8× NVIDIA H200 GPU node — the standard building block for on-prem AI training and heavy inference. It delivers 1,128 GB of HBM3e (8…