Skip to content
Make in India OEM · INR-transparent · Pan-India onsite SLATalk to sales: +91 720 794 8743Sign in

Token Economics for Rack-Scale AI: Cost per Million Tokens, Honestly

Concept Updated 15 Jul 2026 · 5 min read

Overview

Rack-scale AI systems are increasingly sold on token economics — cost per million tokens generated and throughput per megawatt — rather than on FLOPS. NVIDIA’s marketing puts the GB300 NVL72 at figures like 35× lower cost per token than Hopper, with independent-methodology benchmarks (SemiAnalysis InferenceMAX-class) supplying the underlying curves. For a CFO or infrastructure head evaluating a DRACO-class purchase, the vendor multipliers matter less than the method: how to convert one rack into a defensible ₹-per-million-token number for your workload, and how that compares against API pricing and cloud rental. This article builds that model and flags where the marketing math flatters.

Token Economics for Rack-Scale AI: Cost per Million Tokens, Honestly
What you’ll learn: How to compute owned cost-per-million-tokens from first principles, why throughput-per-megawatt became the binding metric, where vendor multipliers come from, break-even against API and cloud pricing, and a worked sensitivity table.

Key takeaways

  • Owned token cost = (amortised capex + power + facility + ops) ÷ tokens actually served — utilisation sits in the denominator and dominates everything.
  • Vendor claims (35× vs Hopper, cents per million tokens) assume high-batch, FP4-optimised serving at near-full utilisation; enterprise duty cycles rarely match.
  • Throughput per megawatt is now the binding constraint in power-limited facilities — and the honest basis for comparing generations.
  • At sustained enterprise utilisation (40–60%), owned rack-scale serving typically undercuts per-token API pricing for open-weight models by a wide margin — but only after the workload exists.
  • Token revenue claims (15× ROI) model API resale economics, not enterprise cost avoidance; use your own value-per-token, not theirs.

The model: five lines, one denominator

Numerator, per month: hardware amortisation (system price over 48–60 months), power (rack IT load × PUE × tariff — a ~140 kW rack at PUE 1.3 and ₹8/kWh runs ~₹10–11 lakh/month), facility (liquid-cooled colo position or amortised fit-out), operations (staff share, service contracts), and network/storage overheads. Denominator: tokens served per month = benchmarked tokens/second for your model-precision-context mix × seconds × utilisation. The utilisation term is where business cases die: a rack benchmarked at millions of tokens/second earns them only when demand fills it. Compute the number at 30%, 50% and 70% — not at the benchmark’s 95%.

Why per-megawatt became the metric

Power, not capital, caps AI deployment in most 2026 facilities — Indian colos allocate liquid-cooled kW before they allocate floor. Measuring tokens per second per megawatt normalises generational comparisons to the actual scarce resource: vendor materials citing large throughput-per-megawatt gains over Hopper reflect real FP4, NVLink-domain and disaggregated-serving advances, even when specific multipliers are best-case. The planning consequence: if your facility ceiling is, say, 300 kW, generation choice determines your token capacity more than rack count — which is also the strongest rational argument for buying the newest platform your schedule supports, per our Blackwell-vs-Rubin timing analysis.

Reading vendor multipliers without being had

Three systematic flatteries. Baseline choice: “35× lower than Hopper” compares against a two-generation-old platform at its worst workload fit (long-context reasoning), not against the B200 fleet you might otherwise buy. Workload fit: multipliers peak on high-batch, FP4-quantised, long-reasoning serving that exploits the 72-GPU NVLink domain; short-context chat on a 70B model shows far smaller deltas — and may not need rack scale at all, per rack-scale vs scale-out. Utilisation assumption: headline cents-per-million-token figures presume the rack is saturated. None of this makes the platform bad — it makes your own benchmark, on your model mix, the only number worth signing against.

Break-even against APIs and cloud rental

Directional 2026 arithmetic: frontier-model APIs bill on the order of $1–15 per million output tokens depending on model class; open-weight serving via APIs runs $0.1–1. An owned NVL72-class rack serving open-weight models at 50% utilisation typically lands in the low cents per million tokens on optimised stacks — one to two orders of magnitude under API list — before counting the sovereignty and DPDP benefits of in-country serving. The catches: capex is committed before demand materialises; engineering an optimised serving stack (TensorRT-LLM/vLLM, disaggregation, quantisation) is real work; and below sustained mid-double-digit utilisation, cloud rental at $2–3/GPU-hour or IndiaAI-subsidised capacity remains the cheaper bridge. The sequencing in our inference sizing checklist — prove demand rented, then buy the baseline — applies with extra force at rack scale.

Sensitivity: what moves the token cost

Lever Swing Effect on ₹/M tokens Owner
Utilisation 30% → 70% 2.3× more tokens −57% Demand planning, consolidation
Precision FP8 → FP4 (where quality holds) ~1.5–2× throughput −33–50% Serving engineering
PUE 1.6 → 1.25 −22% power bill −5–8% Facility choice
Amortisation 36 → 60 months −40% monthly capex line −20–30% Finance policy + mid-life role plan
Batch/disaggregation tuning 1.5–3× throughput −33–66% Serving engineering

Frequently asked questions

What token cost should an owned rack achieve?

Serving open-weight models on an optimised stack at 50%+ utilisation, expect low single-digit cents per million tokens for mid-size models and tens of cents for the largest — but derive it from your own benchmark, tariff and utilisation, not from marketing slides.

Are the 15×-ROI and 35×-cheaper claims false?

Not false — conditional. They model saturated racks, FP4-optimised reasoning workloads and (for ROI) API-resale revenue. Enterprises avoiding API spend at partial utilisation see large but smaller advantages.

Why do power metrics matter more than price per GPU?

Because liquid-cooled kilowatts are the scarce resource in 2026 facilities, including India’s. Tokens per megawatt determines what a fixed facility ceiling can earn — and newer generations win primarily on that axis.

When does API spend justify buying a rack?

A steady seven-figure-rupee monthly API bill for open-weight-servable workloads is the usual tripwire — at that run-rate, an owned or hosted rack with a competent serving team typically pays back within its first two years. Verify with the five-line model at conservative utilisation.

Does this analysis apply to fine-tuning and training?

Partially — training economics are about time-to-result, not tokens served, and favour rental for bursts. The token model governs the serving estate, which is where rack-scale systems spend most of their life.

Ready to deploy?

Talk to an RDP architect about power, cooling and lead time.

Request a Quote
👋 Ask GPU Mart AI — voice & text