Deployment & Commissioning
Site readiness, burn-in, acceptance testing and handover.
9 articlesThe First 90 Days of an AI Cluster: Early-Life Failures, Spares and Warranty
What actually fails in an AI cluster's first 90 days - GPUs and HBM, optics, NVMe, PSUs - and how to prepare: a spares kit sized to your scale, warranty…
AI Cluster Handover and Documentation: What to Demand From Your Integrator
The complete handover pack to demand before your integrator leaves site: as-built rack elevations and rail maps, firmware and serial inventories, switch and BIOS configuration exports, runbooks, credentials and licences,…
Storage and Data-Path Validation: Proving Storage Will Not Starve Your GPUs
How to validate storage during AI cluster commissioning: fio and gdsio test recipes, per-GPU read bandwidth targets, checkpoint write burst testing, metadata and small-file checks, and the combined GPU-plus-storage soak…
GPU Health on Day 1: DCGM Diagnostics, XID Errors and the Baseline to Record
How to establish a GPU health baseline on day one of a new cluster: DCGM diagnostic levels and what each catches, the XID error codes that matter, ECC and row-remapping…
Commissioning a Liquid-Cooled AI Rack: Fill, Leak Test and Coolant Sign-Off
The commissioning sequence for a direct-to-chip liquid-cooled AI rack: pre-fill flush, pressure decay and wet leak testing, fill and air purge, CDU setpoints, and the coolant chemistry numbers - pH,…
Site Readiness for AI Rack Delivery: Power, Floor Loading, Access and Staging
A pre-delivery site readiness checklist for heavy AI racks in India: feeder and breaker capacity, floor loading and point loads, door widths and lift limits along the access route, network…
NCCL Tests and Collective Benchmarks: Validating an AI Fabric Before Production
How to use nccl-tests to prove a GPU fabric works: what bus bandwidth means, the numbers a healthy NVLink and 400G RDMA fabric should hit, how to scale tests from…
Burn-In and Validation for GPU Clusters: Catching Failures Before Production
How to burn in a new GPU cluster: which stress tools to run, how long a thermal soak really needs, what infant-mortality failures look like in logs, and the exit…
Acceptance Testing an AI GPU Cluster: What to Verify Before Sign-Off
A practical acceptance-test framework for AI GPU clusters: the four test layers, the HPL and nccl-tests numbers to demand, DCGM health gates, and a sign-off matrix that defines who signs…