AI Cluster Handover and Documentation: What to Demand From Your Integrator
Overview
Handover is where an AI cluster stops being the integrator’s project and becomes your production system – and the difference between a good and bad handover is measured months later, at 2 a.m., when someone needs to know which leaf switch port rail 5 of node 14 lands on. A complete handover pack contains six things: as-built physical documentation (elevations, wiring and rail maps), a firmware and serial inventory, configuration exports for every device, operational runbooks, credentials and licence transfers, and the full test evidence from commissioning. Demand it as a contractual deliverable with payment linked to it, because no integrator writes documentation enthusiastically after the invoice clears.


Key takeaways
- Make the handover pack a named contract deliverable with a checklist appendix – “complete documentation” without an itemised list is unenforceable.
- As-built means what was actually built: rack elevations, port-level wiring and rail maps, and cable labels that match the spreadsheet – verify by spot-checking 10 percent of connections physically.
- The firmware and serial inventory (GPU, VBIOS, BMC, BIOS, NIC, switch, PSU, drive) is your warranty currency: claims, advisories and security patches are all resolved against it.
- Configuration must be exportable and rebuildable: switch configs, BIOS/BMC profiles, OS image or IaC repository – the test is whether you could re-image a replacement node to identical state without the integrator.
- Runbooks should cover the ugly days: full start-up and shutdown order (critical for liquid-cooled racks), single-node swap, and the escalation matrix with names and response times, not just a support email.
As-built physical documentation
The physical pack starts with rack elevations (every RU accounted for, front and rear), the power map (which PSU feeds from which PDU and phase – phase balance matters when you add load later), and above all the network wiring map at port level: every GPU NIC to its leaf port, every rail’s colour and label, storage and management networks, with cable IDs printed on both ends of the physical cables. On rail-optimised fabrics a wrong-rail cable is a performance bug you will chase for weeks – our rail topology guide explains why the mapping is so easy to get subtly wrong. Verification is physical: before signing, trace a random 10 percent of cables against the map, and reject the pack if the error rate is not zero – documentation that is 95 percent right is worse than none, because you will trust it.
Firmware, serial and licence inventory
Demand a machine-readable inventory (CSV or JSON, not PDF) listing per component: model, serial number, firmware/VBIOS/BIOS version, and location (node, slot, rack, RU). This is the artefact every later event resolves against – warranty claims (“which drives are affected by this advisory?”), security response, and the day-1 health baseline it should be stored beside. Alongside it: licence and entitlement transfer – NVIDIA AI Enterprise or vGPU licences, switch OS licences, support portal registrations moved to your organisation’s account, not left registered to the integrator. In India, also collect the import documentation trail (invoices, bill of entry references) – you will need it for insurance, audits and any AMC dispute about what was actually supplied.
Configuration exports and rebuildability
The standard to hold the integrator to: with the pack alone, a competent engineer could rebuild any single device to its production state. That means exported switch configurations (with a note of the OS version they apply to), BIOS and BMC settings profiles (many vendors support export; at minimum a documented settings delta from defaults – NUMA, power profiles and PCIe settings materially affect AI performance), the OS provisioning method (golden image, Ansible/Terraform repository, or cluster manager configuration such as Base Command Manager or xCAT exports), and the scheduler configuration (Slurm/Kubernetes manifests, partition and QoS definitions). If the integrator provisioned from their private tooling, require the rendered artefacts even if they keep the tooling. Test it during acceptance: pick one node, wipe it, and have your engineer re-provision it from the pack while the integrator watches – this single exercise surfaces more missing documentation than any review meeting.
Runbooks, training and the escalation matrix
Minimum runbook set: full cluster start-up and shutdown sequences with ordering and timing (liquid-cooled racks have strict ordering between CDU, facility water and IT load – see liquid-cooled commissioning); emergency power-off consequences and recovery; single-node drain, swap and return-to-service; GPU replacement procedure with firmware-matching steps; and the monitoring guide (what each alert means, which are pageable). Then the escalation matrix: named contacts with phone numbers for the integrator, each OEM, and the AMC provider, with contracted response and resolution times per severity – and who is allowed to raise RMAs. Insist on a working session, not a slide deck: your operators execute a node swap and a controlled shutdown with the integrator supervising. Steady-state practice then belongs to Day-2 operations.
Test evidence, punch list and keeping it alive
The pack closes with evidence: the signed acceptance test results, burn-in logs, storage and fabric benchmark outputs, coolant chemistry baseline, and the punch list of open items with owners and dates. After handover, the pack only stays true if changes flow through it: store it in version control, require the AMC provider to update the inventory on every part swap, and diff the live fabric against the wiring map after any re-cabling. A handover pack that is not maintained becomes archaeology within a year; one that is maintained becomes the single most-consulted document in the first 90 days – which is exactly when early-life failures arrive.
Handover deliverables checklist
| Deliverable | Format | Verification before sign-off |
|---|---|---|
| Rack elevations + power map | Drawings + spreadsheet | Walk one rack end to end against the drawing |
| Wiring / rail map | Port-level spreadsheet + labels on cables | Physically trace 10% of links; zero errors |
| Firmware + serial inventory | CSV/JSON, machine-readable | Spot-check against nvidia-smi/BMC on sample nodes |
| Config exports (switch/BIOS/BMC/OS/scheduler) | Text configs + repo or images | Wipe-and-rebuild one node from pack alone |
| Runbooks + escalation matrix | Living documents (wiki/repo) | Operators execute node swap + shutdown drill |
| Credentials + licences | Handed over in your secrets manager | All default passwords rotated; portals transferred |
| Test evidence pack | Raw logs + signed certificates | Matches ATP; baselines archived with inventory |
| Punch list | Tracked list with owners/dates | Payment retention mapped to closure |
Frequently asked questions
When should the handover pack be agreed?
At contract time, as an itemised appendix – before hardware ships. Agreeing the list after installation puts you in a weak position: the integrator’s team has moved to the next project and every document becomes a favour. Link 10-20 percent of payment to pack acceptance, and define acceptance as verified, not delivered.
What if the integrator says their tooling is proprietary?
Reasonable for the tooling, unreasonable for the outputs. You are entitled to the rendered artefacts: actual switch configs, actual BIOS settings, a bootable image or documented build for the OS layer. If they cannot hand over enough for an independent engineer to rebuild a node, you do not own your cluster – you rent your integrator.
Who should hold the credentials at handover?
All credentials transfer into your secrets management at handover, every default and integrator-set password is rotated in your presence, and the integrator’s remote access is removed or re-issued under your control with time-limited, audited accounts if an AMC requires it. “We keep admin for support convenience” is a security finding, not a service.
How detailed should the wiring map be for a small cluster?
The same port-level standard, just smaller – a 4-node cluster’s map is an afternoon’s work. Small clusters are actually where maps go missing most, because everyone assumes they can remember 100 cables. They cannot, and the person who cabled it will not be the person debugging it.
What documentation applies if we bought a pre-integrated rack?
Factory-integrated racks (with integration done off-site) arrive with the vendor’s build documentation – but you still need the site-specific layer: facility connections, uplinks into your network, management integration, and everything in the runbook and escalation sections. Ask for the factory test reports too; they are the pre-shipment half of your acceptance evidence.
Ready to deploy?
Talk to an RDP architect about power, cooling and lead time.