

DRACO 16× H200 SXM Rack-Scale AI System
2-node HGX H200 · 4× Intel Xeon 6 · 4 TB DDR5 ECC · 120 TB NVMe · Quarter-Rack · liquid-cooled
Key Specifications
See full specs ↓“RDP delivered and installed our edge AI pods across 6 sites with predictable INR pricing and onsite SLA.” — [customer / sector, to confirm]


Designed, built and supported in India — sovereign by design
Your AI factory on sovereign Indian infrastructure: data residency under DPDP, MeitY-recognised, ISO 27001 / SOC 2 deployment paths, and procurement on GeM.
Overview
The DRACO 16× H200 SXM Rack-Scale AI System is a turnkey, liquid-cooled quarter-rack AI training cluster — 16 NVIDIA H200 SXM GPUs across 2 HGX nodes, wired into one system with NVLink inside each node and a 400G InfiniBand spine between them. It delivers 2,256 GB HBM3e of aggregate HBM3e and arrives racked, cabled, cooled and tested, ready to train and serve the largest models on-premises — in INR, on a GST invoice.
Engineered for organisations building serious in-house AI capacity, it removes the integration risk of assembling a cluster yourself: RDP sizes the nodes, fabric, storage, power and cooling as one validated system and delivers it as a single SKU with one warranty and one support contract.
Key highlights
- 16× H200 SXM · 2,256 GB HBM3e aggregate HBM3e — cluster-scale GPU memory for trillion-token training and large-model serving.
- NVLink + NVSwitch in each node, 400G InfiniBand spine — full intra-node bandwidth and low-latency inter-node collectives for near-linear scaling.
- 4× Intel Xeon 6 + 4 TB DDR5 ECC — host compute and memory matched to 16 data-centre GPUs.
- 120 TB NVMe NVMe + parallel-FS ready — high-throughput data and checkpoint storage across the cluster.
- Quarter-Rack, liquid-cooled, turnkey — delivered racked, cabled, cooled and burned-in; one SKU, one warranty.
- On-prem data sovereignty — training data and weights stay in-house; DPDP-friendly, air-gappable.
- Make-in-India OEM — predictable INR pricing, GST tax invoice (HSN 8471), pan-India onsite support, GeM-procurable.
- Scale path — grow to larger rack-scale systems and multi-rack AI SuperClusters on the same fabric.
AI workload fit (what it actually runs — honestly)
- Distributed training: data-, tensor- and pipeline-parallel training of 405B+-class models across the 16 GPUs.
- Large-scale inference: serve many large models, or shard the largest models across nodes for high throughput.
- RAG, multimodal & agentic at scale: production AI platforms on the cluster’s storage and fabric.
- Engineering note: within each HGX node the eight GPUs share an NVLink+NVSwitch fabric; across nodes, a 400G InfiniBand spine carries the collective (all-reduce) traffic. Real-world scaling efficiency depends on model and parallelism strategy — we validate it for your workload rather than quoting a peak FLOPS number.
AI workload positioning
This sits at the cluster-scale train-and-serve stage: a complete, validated AI system rather than a single server. With 2,256 GB HBM3e of HBM3e and a balanced NVLink/InfiniBand fabric, it is sized to sustain real distributed training and large-scale serving — the owned alternative to renting an equivalent cloud cluster, where the meter never stops.
Industry use cases
- Government & sovereign AI — a national or departmental AI cluster on GeM-procurable infrastructure.
- BFSI — a private training cluster for large risk, fraud and language models.
- Healthcare & life sciences — large-scale model and imaging training, data in-house.
- Neocloud / AI providers — a validated rack to build or expand a GPU cloud.
- Research & higher-ed — an institutional AI training cluster.
- Large enterprise & telecom — in-house foundation-model and platform development.
Performance — and how to be sure
We don’t publish inflated peak numbers. The honest picture: 2,256 GB HBM3e of HBM3e across 16 NVSwitch-linked GPUs on a 400G InfiniBand spine is sized for 405B+-class distributed training and large-scale serving. Want certainty? Request a free benchmark of your model and dataset — including a scaling test across the nodes — on this exact configuration before you buy; we’ll send back real tokens/sec, scaling efficiency and timings.
Series & upgrade path
- DRACO (flagship rack-scale tier) — this.
- Scale ladder: 16-GPU quarter-rack → 32-GPU half-rack → 64-GPU full-rack; step up to Blackwell B200/GB200/GB300 rack-scale systems for more memory per GPU.
- When to step up: for multi-rack scale, move to RDP AI SuperClusters built from these systems — talk to an architect about the fabric and facility.
On-prem vs cloud — the TCO case
For a sustained training cluster, owning beats renting decisively: an always-on cloud cluster of this size is the dominant line in an AI budget, and on-prem removes egress fees and keeps data and weights in-house. RDP pricing is fixed in INR with a GST input-credit-eligible invoice — ask for a 3-year cluster TCO comparison.
Software & day-one readiness
Ships pre-integrated and validated: NVIDIA driver, CUDA, cuDNN, NCCL, the InfiniBand stack, Docker and NVIDIA Container Toolkit, with Slurm or Kubernetes, PyTorch and vLLM / Triton / TensorRT-LLM on Ubuntu LTS. Optional managed cluster operations, scheduler and observability setup.
Power, cooling & rack integration
A quarter-rack liquid-cooled system — plan facility power, CDU/manifold and water, and the InfiniBand spine. RDP scopes power, cooling and floor/rack requirements as part of the design. (Exact PSU/PDU ratings, BTU, flow and facility figures confirmed on the build sheet.) Full out-of-band management across the cluster.
Deployment, warranty & support
- Made to order, integrated, racked, cabled, cooled and burned-in in India; realistic lead time confirmed at quote.
- Delivered as one system: nodes, fabric switches, cabling, PDUs, cooling integration, and the pre-installed cluster software stack.
- Onsite warranty + AMC with pan-India coverage, cluster-level support and an RMA/escalation path (exact term & response window confirmed at quote).
Why RDP
14 years of Make-in-India infrastructure and 300,000+ devices shipped. Indian OEM, INR pricing, GST tax invoice (HSN 8471), pan-India onsite engineers, GeM availability, and DPDP / sovereign-AI-ready deployment.
Buy with confidence
This is a turnkey rack-scale AI system, made to order — talk to an RDP solution architect, size the cluster, fabric and facility, get a 3-year TCO, and benchmark your own model with a scaling test before you commit. Request a quote to begin.
Specifications
| Use Case | Agentic AI, Computer Vision, Fine-tuning, Generative AI, HPC & AI, Inference, LLM Training, NLP & Speech, RAG, Sovereign AI |
| GPU Model | NVIDIA H200 SXM5 |
| Form Factor | Rack |
| Workload Fit | Entry rack-scale training |
| GPUs | 2-node (16× NVIDIA H200 SXM5) |
| GPU memory | 2,256 GB HBM3e (16× 141 GB) |
| Model fit | 405B+ |
| CPU | 4× Intel Xeon 6 |
| System memory | 4 TB DDR5 ECC |
| Storage | 120 TB NVMe |
| Networking | NVLink + InfiniBand NDR 400G |
| Chassis | Quarter-Rack |
| GPU Count | 16 |
| Cooling | Liquid |
| Series | DRACO |
Why RDP GPU Mart
- ✓ Make in India OEM — Hyderabad facility, 14 years, 300,000+ devices shipped.
- ✓ Sovereign-ready: India data residency (DPDP), MeitY-recognised, ISO 27001 / SOC 2 paths.
- ✓ INR-transparent: GST invoice, CGST/SGST or IGST, pan-India onsite SLA.
- ✓ Available on GeM for government and PSU procurement.
FAQ
Is GST invoicing available?
Yes — GST invoice, CGST+SGST or IGST by billing state, eligible for input credit.
Do you deliver and install pan-India?
Yes — pan-India delivery with onsite installation and a 3-year onsite SLA.
What warranty and support is included?
3-year pan-India onsite SLA with AMC and flexible financing options.
Can this be configured to my workload?
Yes — talk to an RDP solutions architect for a custom build or multi-node cluster.
Compare the range
Other Rack-Scale AI Systems in this line
Swipe to compare
| GB300 NVL72 Rack-… | 64× H200 SXM Rack… | 32× H200 SXM Rack… | 32× B200 SXM Rack… | |
|---|---|---|---|---|
| GPUs | GB300 NVL72 (72× Grace-Blackwell Ultra) | 8-node (64× NVIDIA H200 SXM5) | 4-node (32× NVIDIA H200 SXM5) | 4-node (32× NVIDIA B200 SXM) |
| GPU memory | 20,736 GB HBM3e (72× 288 GB) | 9,024 GB HBM3e (64× 141 GB) | 4,512 GB HBM3e (32× 141 GB) | 5,760 GB HBM3e (32× 180 GB) |
| Model fit | Trillion-scale frontier | Trillion-scale | Trillion-scale | Trillion-scale |
| Networking | NVLink domain + InfiniBand spine | NVLink + InfiniBand NDR 400G | NVLink + InfiniBand NDR 400G | NVLink + InfiniBand NDR 400G |
| Chassis | Single-Rack (NVLink domain) | Full-Rack | Half-Rack | Half-Rack |
| Price | On request | On request | Request a Quote | Request a Quote |
| Quote | Quote | View | View |
Build the full stack
Pair it with








Designing a GPU cluster, not just one server?
Talk to an RDP solutions architect about the full fabric — networking, storage, rack and power.
*Pan-India delivery and onsite installation are subject to location serviceability; standard SLA terms apply. Specifications indicative; final configuration confirmed on quote.