Skip to content
Make in India OEM · INR-transparent · Pan-India onsite SLATalk to sales: +91 720 794 8743Sign in

Inference

8 articles
15 Jul 2026

Disaggregated Inference: Splitting Prefill and Decode in 2026

Prefill is compute-bound, decode is memory-bound, and 2026 serving stacks split them onto dedicated GPU pools with KV cache shipped between - worth 2-3x throughput at node scale and far…

5 min readRead →
10 Jul 2026

GPU Server India Sizing Checklist for AI Inference

When sizing GPU servers for AI inference in India, consider memory capacity, workload requirements, and compliance with data protection regulations. The NVIDIA H200's 141 GB HBM3e memory and adherence to…

5 min readRead →
8 Jul 2026

Sizing 70B LLM Inference on GPU Servers in India

A 70B LLM inference plan should start with memory, concurrency, latency, power, and data-residency constraints, not only GPU count. For Indian teams, RDP GPU Mart can turn those constraints into…

4 min readRead →
6 Jul 2026

GPU Sizing for Agentic AI Workloads (2026)

Agentic AI changes the sizing question: it isn't just a GPU problem. The GPU still runs the model, but the orchestration layer — scheduling sub-tasks, routing tool calls, passing state…

4 min readRead →
6 Jul 2026

How Many GPUs Do You Need for a 100-User Private ChatGPT (On-Prem)?

A private, on-prem ChatGPT-style assistant for 100 users typically runs on one to two datacenter GPUs. Because "100 users" rarely means 100 people typing at once — active-to-total ratios are…

4 min readRead →
6 Jul 2026

Sizing GPU Compute for Computer Vision / Video Analytics

Computer-vision and video-analytics GPU sizing is throughput-driven, not memory-driven: capacity is set by streams × resolution × frame rate × model complexity, and constrained by the GPU's video-decode and inference…

4 min readRead →
6 Jul 2026

How Much VRAM Does a 70B / 405B LLM Need for Inference?

For inference, a 70B-parameter model needs ~140 GB of GPU memory at FP16, ~70 GB at FP8, and ~43 GB at 4-bit — so one to two datacenter GPUs. A…

4 min readRead →
Reference Architecture 21 Jun 2026

GB300 NVL72: Anatomy of a 120 kW Rack-Scale AI Factory

Overview The NVIDIA GB300 NVL72 (Blackwell Ultra) marks the point where the rack, not the GPU, becomes the unit of compute. Seventy-two Blackwell Ultra (B300) GPUs and 36 Grace CPUs…

4 min readRead →

Need help in Inference?

Request a Quote