Products & Installation
48 articlesPCIe Gen6 and 800G SuperNICs: What the 2026 Server Refresh Changes
ConnectX-8 combines an 800G NIC with an integrated PCIe Gen6 switch in one device, exposing 48 Gen6 lanes from an x16 connector. That consolidation removes components from AI server designs…
Workstation Thermals and Power in Indian Offices: Planning for 600 W Cards
A modern AI workstation can draw over a kilowatt at the wall and runs at that level for hours, not seconds. In an Indian office that is a thermal and…
Local RAG on a Mid-Tier Workstation: Index Size and Honest Limits
A mid-tier workstation can host a genuinely useful private RAG system over a few hundred thousand documents. This sets out where the memory goes between generation model, embeddings and index,…
The Multi-Agent Developer Desk: Sizing a Mid-Tier Workstation
Developers now run several concurrent agent sessions rather than one chat. Concurrency multiplies KV cache rather than weights, which is why a 96 GB mid-tier workstation behaves very differently from…
Scale-Up Domains and the Fabric Boundary: Designing Past One Rack
Every rack-scale system has a boundary where fast NVLink ends and slower scale-out networking begins. Where you place that boundary relative to your model's parallelism determines achieved throughput more than…
Liquid Cooling for Flagship Racks in India: CDUs and Water Quality
Direct liquid cooling is mandatory above roughly 50 kW per rack, and the part Indian operators underestimate is water chemistry: deionised or RO water with conductivity typically held below 5-10…
Rubin Ultra and Kyber NVL576: Planning the 2027 Flagship Rack
Kyber is expected to house 576 Rubin Ultra GPUs per rack at 600 kW to 1 MW, with 800 VDC distribution arriving alongside it in 2027. Any hall commissioned in…
Entry Workstation Refresh Cycles: When to Replace and When to Wait
The refresh question in 2026 is not raw speed but VRAM and FP4 support. A card that cannot hold the models your team now uses is obsolete regardless of its…
Running Local Agents on Entry Hardware: Limits and Workarounds
Agentic workflows multiply model calls and grow context with every step, which punishes small VRAM budgets harder than chat does. This sets out what genuinely runs on a 24-32 GB…
Entry AI Workstations for Indian Colleges and Research Labs (2026)
Academic AI labs in 2026 face a specific squeeze: memory prices up sharply, curricula demanding hands-on GPU work, and subsidised national compute available for bursts. This sets out how to…
Vision AI on a Single 32 GB Card: What an Entry Workstation Handles
A 32 GB Blackwell-class workstation card runs detection, segmentation, OCR and mid-size vision-language models comfortably. This sets out realistic throughput expectations, where 32 GB stops being enough, and how to…
Desk-Side Mini AI Boxes in 2026: Where Unified Memory Fits
Mini AI boxes with large unified memory pools load models a discrete GPU cannot hold, but at a fraction of the memory bandwidth. That trade decides what they are good…
The 2026 Memory and NAND Squeeze: Procuring AI Storage Under Allocation
NAND contract prices were reported rising 70-75 percent quarter-on-quarter in Q2 2026 as fabs shifted capacity to HBM. With new fab output unlikely before late 2027, AI storage procurement in…
Checkpoint Storage at Frontier Scale: Sizing for 2026-2028
Published work puts checkpoint overhead at 12-43 percent of total training time, with a 16,000-accelerator cluster taking roughly 155 checkpoints a day. This explains how to size checkpoint capacity, bandwidth…
Vector Index Sizing: From 10 Million to 1 Billion Embeddings
A billion 1536-dimension vectors in a plain HNSW index needs terabytes of RAM. Quantization cuts that by 4x to 32x, and disk-based indexes move the graph to NVMe. This is…
GPUDirect Storage and DPU Offload: Designing the AI Data Path
GPUDirect Storage moves data between NVMe and GPU memory without a host bounce buffer; DPU offload removes the storage host from the path entirely. Together they define the 2026 AI…
Object Storage Enters the Training Loop: S3-over-RDMA in 2026
Object storage used to be the cold tier behind a parallel filesystem. In 2026, S3-over-RDMA and DPU-resident object stores put it directly in the GPU data path, which changes how…
The Air-Cooled Middle: RTX PRO Servers Between Workstation and HGX
MGX-based RTX PRO servers pack up to eight 96 GB GDDR7 GPUs - 768 GB aggregate - into standard air-cooled racks, serving multiple 70B-class replicas plus rendering and vGPU duty…
KV Cache Offloading: The New Storage Tier in AI Inference Servers
Long-context and agentic inference made the KV cache a storage problem: about 300-350 KB per token on 70B-class models spills from VRAM to DRAM, NVMe and shared tiers, with LMCache-class…
Self-Hosted Coding Assistants: Running Coder LLMs on Entry Workstations
Open 24-32B coder models now benchmark near commercial APIs and run on a 24-32 GB entry workstation with 30-60 ms local completion latency cloud endpoints cannot match. For Indian services…
Entry AI Workstations in 2026: What 16, 24 and 32 GB Really Run
At the entry tier, VRAM decides everything: 16 GB is a learning machine, 24 GB a value point for 7-13B work, and 32 GB comfortably serves quantised 30B models and…
Day-2 Operations for Rack-Scale AI: Failures, Monitoring and Service
Dense GPU systems fail as routine - Meta logged 419 interruptions in 54 days at 16k-GPU scale - so rack-scale readiness means DCGM-based trend monitoring, 30-60 minute checkpoint cadence, trained…
Token Economics for Rack-Scale AI: Cost per Million Tokens, Honestly
Owned token cost is amortised capex, power, facility and ops divided by tokens actually served - utilisation dominates. Vendor multipliers like 35x vs Hopper assume saturated FP4 reasoning workloads; benchmark…
Sovereign AI Pods: Rack-Scale Planning for India-Controlled Compute
A sovereign AI pod is 1-4 racks of compute, storage and fabric operated under Indian jurisdiction: local keys, cleared admin access, contained telemetry and in-country logs. IndiaAI public capacity near…
Buy Blackwell Ultra or Wait for Rubin? Flagship Timing for 2026-27
Vera Rubin is in production with cloud shipments from H2 2026, but enterprise racks realistically land in 2027. Deployed GB300-class output for 12-18 months usually beats the successor delta; wait…
One Rack or Nine Nodes: The Rack-Scale vs Scale-Out Decision
The flagship-tier choice is unit of scale: nine HGX B300-class nodes or one GB300 NVL72 rack fusing 72 GPUs and about 21 TB of HBM into a single 130 TB/s…
Hosting 100 kW Racks in India: Facility Readiness for Rack-Scale AI
A 120 kW NVL72-class rack exceeds most legacy Indian hall designs, so facility readiness is the long pole: direct-to-chip liquid cooling, 415 V high-amperage feeds, 2-tonne floor loading and contracted…
AI Workstation TCO in India: Duties, GST and the Cloud Crossover
GPU hardware enters India at 0% basic duty under ITA-1, and the 18% IGST is input-creditable, so the buy-vs-rent question is pure utilisation arithmetic. Against $2-3/GPU-hour cloud rates, a daily-driver…
From Desk to Server Room: When a Team Outgrows AI Workstations
Four signals say a team has outgrown workstations: GPU queueing, duplicated model weights, uptime needs and VRAM ceilings. Measure two weeks of utilisation, then buy a boring first server -…
One Card, Two Pipelines: Hybrid Rendering and AI Workstations in 2026
Rendering and AI converged on the same silicon: 96 GB Blackwell workstation cards run V-Ray by day and diffusion or 70B inference overnight, while DLSS 4 and neural texture compression…
Fine-Tuning LLMs on a Workstation: LoRA and QLoRA Memory Math
Fine-tuning memory is weights plus gradients, optimiser states and activations. QLoRA needs about 12 GB for 7B, 44 GB for 32B and 88 GB for 70B, so a 96 GB…
Dual-GPU AI Workstations Without NVLink: Planning a 192 GB Tower
Workstation Blackwell cards have no NVLink, so a dual 96 GB tower pools 192 GB over PCIe 5.0. Choose 300 W Max-Q cards for multi-GPU builds, favour pipeline parallelism and…
The 96 GB Desk-Side Tier: What a Mid-Range AI Workstation Runs in 2026
A 96 GB-class workstation card (RTX PRO 6000 Blackwell: 24,064 CUDA cores, ECC GDDR7 at ~1.8 TB/s) runs quantised 70B inference and 7-32B fine-tuning on a desk. The ceiling: unquantised…
Small Language Models and the NPU Myth: What Actually Runs Locally in 2026
A 3–9B model now carries most of an agentic loop locally, faster and more privately than a cloud API. But the NPU is not what runs it: Ollama, llama.cpp and…
Local AI Workstations in 2026: Running 70B Models at Your Desk
With 96 GB of GDDR7 on a single RTX PRO 6000 Blackwell card, a desk-side workstation can now hold a 70B model at Q8 — and two Max-Q cards pool…
GPU Storage Planning for LLM Checkpoints and RAG Indexes
LLM checkpoint and RAG index storage demands are determined by model size, checkpoint frequency, and retrieval corpus scale. A 70B-parameter model can generate checkpoints exceeding 140 GB per save, while…
NVIDIA GB300 NVL72 Supercluster: Inside the 8-Rack Containerised AI Factory Node
The NVIDIA GB300 NVL72 supercluster packs eight NVL72 racks — 576 Blackwell Ultra B300 GPUs and 288 Grace CPUs — into one containerised AI factory node delivering ~11.5 EFLOPS FP4,…
AI Workstation vs GPU Server: RDP Buyer Decision Guide
Choosing between an AI workstation and a GPU server involves understanding performance needs, workload types, and governance requirements. Workstations are typically suited for individual tasks, while GPU servers offer scalability…
GPU RDP Workstation Buyers Guide for Indian AI Teams
When selecting a GPU RDP workstation for AI teams in India, consider the balance between performance, memory capacity, and compliance with data governance regulations. The NVIDIA H200 and H100 GPUs…
AI Factory Storage Planning for GPU Clusters
AI storage planning should begin with dataset movement, checkpoint frequency, metadata behavior, and recovery objectives before selecting capacity. A GPU cluster that starves on I/O wastes accelerator budget. RDP GPU…
GPU Workstations vs Servers for AI Teams in India
Choose a GPU workstation when one team needs local iteration, controlled data access, and fast developer feedback. Choose a GPU server when concurrency, shared scheduling, larger models, stronger uptime expectations,…
H100 vs H200 vs B200: Which GPU for Your Workload?
Pick by memory and workload: the H100 (80 GB, 3.35 TB/s) is the proven workhorse for training and inference where 80 GB fits; the H200 (141 GB HBM3e, 4.8 TB/s)…
Storage Architecture for AI Training: Why the Bottleneck Isn’t the GPU
In large AI training, the most common bottleneck isn't GPU compute — it's storage failing to feed the GPUs fast enough. Slow storage leaves expensive GPUs idle waiting on data…
GPU Workstation or GPU Server? A Decision Guide
Choose a GPU workstation when one person needs GPU power at their desk for development, fine-tuning experiments, or content creation — typically 1–4 GPUs, deskside, single-user. Choose a GPU server…
InfiniBand vs Spectrum-X vs Ethernet for AI Clusters
Choose the fabric by workload: InfiniBand for training and HPC, where GPU-to-GPU latency dominates all-reduce; Spectrum-X or RoCE Ethernet for inference and multi-tenant clouds, where cost and interoperability matter more…
Air-Cooled vs Liquid-Cooled GPU Racks: When to Switch
Air cooling runs out of headroom at roughly 35 kW per rack. Below that, well-designed airflow is fine; above it, direct-to-chip liquid cooling becomes necessary, and beyond ~100 kW per…
What Is an AI Factory? Rack-Scale AI Explained for Buyers
An AI factory is a rack-scale computing system engineered to do one thing at industrial scale: turn electricity and data into AI output (tokens). Instead of a single server, it…
GB300 NVL72: Anatomy of a 120 kW Rack-Scale AI Factory
Overview The NVIDIA GB300 NVL72 (Blackwell Ultra) marks the point where the rack, not the GPU, becomes the unit of compute. Seventy-two Blackwell Ultra (B300) GPUs and 36 Grace CPUs…