Reference Architectures
18 articlesStorage Network Design for AI Clusters: Dedicated vs Converged Fabric
Should storage traffic share the GPU compute fabric or get its own network? A reference-architecture view of the three fabrics in an AI cluster, when convergence is safe, when checkpoint…
Network and Access Security for GPU Clusters: Segmentation, Jump Hosts and Secrets
A GPU cluster concentrates an organisation's most valuable data, most expensive compute and most privileged credentials in one place. A reference security architecture: network zones, RDMA fabric exposure, jump-host access…
Data Residency Architectures for AI: Where Personal Data May and May Not Flow
DPDP permits cross-border transfer except to restricted countries, so real residency duties come from sector overlays - RBI payment data, government workloads, contracts. Three reference architectures for keeping personal data…
Open Rack v3 (ORv3) Explained: Busbars, Power Shelves and 21-Inch Trays
ORv3 replaces per-server AC cords with a rack-level 48 V DC busbar fed by centralised power shelves, and adds blind-mate connections plus 21-inch tray support. What changes mechanically and electrically,…
Rubin Ultra and Kyber NVL576: Planning the 2027 Flagship Rack
Kyber is expected to house 576 Rubin Ultra GPUs per rack at 600 kW to 1 MW, with 800 VDC distribution arriving alongside it in 2027. Any hall commissioned in…
Federated Learning Across Hospitals: GPU and Network Planning
Federated learning trains a shared model without pooling patient data, which is why it appeals to Indian hospital networks under DPDP. The infrastructure cost is real: every participating site needs…
Engineering Copilots and PLM: On-Prem GPU Planning for Design Data
Engineering knowledge sits in CAD, PLM records, drawings and standards, not in prose. Building a copilot over it is a multimodal retrieval problem with a hard confidentiality constraint, which is…
Physical AI in Indian Factories: GPU Planning for Robotics Pilots
Physical AI needs three distinct compute tiers: simulation for training policies, a training cluster for the models, and edge inference on the robot. Indian manufacturers including Ola Electric and Wipro…
Agentic AI in Banking: GPU Infrastructure Under FREE-AI
Agentic AI in banking multiplies inference per business action and adds an audit obligation for every step. Under RBI's FREE-AI framework, that combination pushes Indian banks toward owned, in-country GPU…
GPUDirect Storage and DPU Offload: Designing the AI Data Path
GPUDirect Storage moves data between NVMe and GPU memory without a host bounce buffer; DPU offload removes the storage host from the path entirely. Together they define the 2026 AI…
Vera Rubin NVL144: What the 2026 Training Platform Changes for Cluster Design
Vera Rubin NVL144 keeps the rack as the scale-up domain but raises memory, interconnect and power together. HBM4, NVLink 6 and ConnectX-9 change how many racks a training run needs,…
RAG GPU Server Reference Architecture for India
A RAG GPU server for India needs fast vector retrieval, low-latency LLM inference, and local data residency in a single coherent architecture. The right design balances HBM-class GPU memory for…
NVIDIA GB300 NVL72 Supercluster: Inside the 8-Rack Containerised AI Factory Node
The NVIDIA GB300 NVL72 supercluster packs eight NVL72 racks — 576 Blackwell Ultra B300 GPUs and 288 Grace CPUs — into one containerised AI factory node delivering ~11.5 EFLOPS FP4,…
Reference Architecture for RAG on H200 GPU Servers
A RAG stack is a retrieval, storage, inference, and governance system, not just a vector database attached to a model. The practical design uses H200-class GPU servers for generation, CPU/storage…
Storage Architecture for AI Training: Why the Bottleneck Isn’t the GPU
In large AI training, the most common bottleneck isn't GPU compute — it's storage failing to feed the GPUs fast enough. Slow storage leaves expensive GPUs idle waiting on data…
Reference Architecture: Sovereign AI Cluster (Scalable Unit)
This reference architecture specifies an in-country, DPDP-aware GPU cluster built from a repeatable Scalable Unit (SU): a group of 8× H200 nodes joined by NDR/XDR InfiniBand, with shared parallel storage…
Reference Architecture: 8× H200 On-Prem AI Training Node
This reference architecture specifies a single 8× NVIDIA H200 GPU node — the standard building block for on-prem AI training and heavy inference. It delivers 1,128 GB of HBM3e (8…
GB300 NVL72: Anatomy of a 120 kW Rack-Scale AI Factory
Overview The NVIDIA GB300 NVL72 (Blackwell Ultra) marks the point where the rack, not the GPU, becomes the unit of compute. Seventy-two Blackwell Ultra (B300) GPUs and 36 Grace CPUs…