Skip to content
Make in India OEM · INR-transparent · Pan-India onsite SLATalk to sales: +91 720 794 8743Sign in

Reference Architecture: Sovereign AI Cluster (Scalable Unit)

Updated 6 Jul 2026 · 4 min read

This reference architecture specifies an in-country, DPDP-aware GPU cluster built from a repeatable Scalable Unit (SU): a group of 8× H200 nodes joined by NDR/XDR InfiniBand, with shared parallel storage and management. You validate one SU, then replicate it to grow capacity predictably — the standard way to build sovereign AI infrastructure that stays under national control. The design keeps data resident, compute domestic, and support local.

Reference Architecture: Sovereign AI Cluster (Scalable Unit)

TL;DR — what's validated

  • Unit of scale = the Scalable Unit (SU): N × 8-GPU H200 nodes + fabric + storage; replicate to grow.
  • Fabric: NDR InfiniBand (400 Gbps/dir) now, XDR 800G (Quantum-X800) for large fabrics — with in-network SHARP reduction (NVIDIA Quantum-X800).
  • Sovereignty by design: in-country siting, DPDP data residency, India-based support.
  • Grow by replication: validate 1 SU → 2 → N; capacity, power, and network scale in known steps.

Overview & the building block

Sovereign clusters are built bottom-up from a node (8× H200, 1,128 GB HBM3e), grouped into a Scalable Unit (several nodes on a shared fabric + storage), then replicated into a cluster. Designing in SUs makes power, cooling, and network growth predictable and auditable — important when a cluster serves national or regulated workloads.

Key components

Compute. 8× H200 nodes (see the 8× H200 node RA) as the atom; SU size (e.g. 4–16 nodes) is set by the largest training run you must fit in one low-latency domain.

Network fabrics. NDR InfiniBand (400 Gbps/dir, ConnectX-7) delivers ~350 GB/s effective all-reduce on an 8-node H100-class cluster with sub-microsecond latency; SHARP performs reductions in the switch, cutting all-reduce round-trips for large clusters (Spheron, 2026). For very large sovereign fabrics, Quantum-X800 XDR 800G (144×800G, SHARP v4) doubles per-port bandwidth (NVIDIA). Separate OOB management plane.

Storage. Shared all-flash / parallel file system sized to keep every SU's GPUs fed; checkpoint bandwidth scales with node count.

Software/management. Cluster scheduler, NCCL/collectives, ZeRO/FSDP for cross-node sharding, monitoring, and multi-tenant isolation if the cluster is shared.

Design requirements

  • Power/cooling: ~10 kW/node; consolidated SUs cross into liquid cooling past ~35 kW/rack.
  • Data residency: all storage and compute in-country; access governed under DPDP.
  • Redundancy: N+1 at the SU level; fabric and storage designed for node loss.

Validated configurations

  • 1 SU (e.g. 4× nodes = 32 GPUs): departmental / single-tenant sovereign workloads.
  • Multiple SUs: national/enterprise scale; validated by replication, not by stretching one SU indefinitely.
  • Larger is supported but customer-specific — validated to the SU; cluster-wide topology is designed to the workload.

Assumptions & scope

H200-node SUs on InfiniBand; figures are planning specs — validate fabric, storage, and facility with your site. Sovereignty depends on in-country siting + DPDP governance, not just hardware. Specs current July 2026.

Where RDP GPU Mart fits

RDP GPU Mart builds sovereign clusters from DRACO nodes and Scalable Units — India-designed, manufactured, and supported, INR-transparent, and DPDP-aware end-to-end, so a PSU or enterprise can stand up in-country AI capacity that stays under its control. *(Request a sovereign-AI cluster design or quote at RDP GPU Mart.)*

FAQ

What is a Scalable Unit in an AI cluster? A repeatable building block — several 8-GPU nodes on a shared fabric and storage — that you validate once and replicate to grow capacity predictably.

Which fabric for a sovereign training cluster? NDR InfiniBand today, XDR 800G (Quantum-X800) for very large fabrics — both with in-network SHARP reduction (NVIDIA).

How do I keep a cluster sovereign? Site compute and storage in-country, govern data under DPDP, and use India-based hardware and support.

How do I scale capacity? Replicate the Scalable Unit — validate one, then add more — rather than stretching a single unit beyond its tested size.

Related

  • Reference Architecture: 8× H200 On-Prem AI Training Node
  • InfiniBand vs Spectrum-X vs Ethernet for AI Clusters
  • Sovereign AI in India: Building In-Country GPU Infrastructure

Research log (Rule #1)

1. NVIDIA — Quantum-X800 InfiniBand (144×800G, SHARP v4). https://www.nvidia.com/en-us/networking/products/infiniband/quantum-x800/ 2. Spheron (2026) — GPU networking IB vs RoCE vs Spectrum-X (all-reduce, SHARP). https://www.spheron.network/blog/gpu-networking-infiniband-roce-spectrum-x-guide/ 3. NVIDIA — H200 datasheet (node memory). https://www.nvidia.com/en-us/data-center/h200/ 4. Network World (2026) — rack density → liquid cooling. https://www.networkworld.com/article/4149069/why-ai-rack-densities-make-liquid-cooling-nonnegotiable.html 5. IndiaAI (2026) — sovereign compute context. https://indiaai.gov.in/hub/indiaai-compute-capacity

Ready to deploy?

Talk to an RDP architect about power, cooling and lead time.

Request a Quote