Token Economics for Rack-Scale AI: Cost per Million Tokens, Honestly
Overview
Rack-scale AI systems are increasingly sold on token economics — cost per million tokens generated and throughput per megawatt — rather than on FLOPS. NVIDIA’s marketing puts the GB300 NVL72 at figures like 35× lower cost per token than Hopper, with independent-methodology benchmarks (SemiAnalysis InferenceMAX-class) supplying the underlying curves. For a CFO or infrastructure head evaluating a DRACO-class purchase, the vendor multipliers matter less than the method: how to convert one rack into a defensible ₹-per-million-token number for your workload, and how that compares against API pricing and cloud rental. This article builds that model and flags where the marketing math flatters.


Key takeaways
- Owned token cost = (amortised capex + power + facility + ops) ÷ tokens actually served — utilisation sits in the denominator and dominates everything.
- Vendor claims (35× vs Hopper, cents per million tokens) assume high-batch, FP4-optimised serving at near-full utilisation; enterprise duty cycles rarely match.
- Throughput per megawatt is now the binding constraint in power-limited facilities — and the honest basis for comparing generations.
- At sustained enterprise utilisation (40–60%), owned rack-scale serving typically undercuts per-token API pricing for open-weight models by a wide margin — but only after the workload exists.
- Token revenue claims (15× ROI) model API resale economics, not enterprise cost avoidance; use your own value-per-token, not theirs.
The model: five lines, one denominator
Numerator, per month: hardware amortisation (system price over 48–60 months), power (rack IT load × PUE × tariff — a ~140 kW rack at PUE 1.3 and ₹8/kWh runs ~₹10–11 lakh/month), facility (liquid-cooled colo position or amortised fit-out), operations (staff share, service contracts), and network/storage overheads. Denominator: tokens served per month = benchmarked tokens/second for your model-precision-context mix × seconds × utilisation. The utilisation term is where business cases die: a rack benchmarked at millions of tokens/second earns them only when demand fills it. Compute the number at 30%, 50% and 70% — not at the benchmark’s 95%.
Why per-megawatt became the metric
Power, not capital, caps AI deployment in most 2026 facilities — Indian colos allocate liquid-cooled kW before they allocate floor. Measuring tokens per second per megawatt normalises generational comparisons to the actual scarce resource: vendor materials citing large throughput-per-megawatt gains over Hopper reflect real FP4, NVLink-domain and disaggregated-serving advances, even when specific multipliers are best-case. The planning consequence: if your facility ceiling is, say, 300 kW, generation choice determines your token capacity more than rack count — which is also the strongest rational argument for buying the newest platform your schedule supports, per our Blackwell-vs-Rubin timing analysis.
Reading vendor multipliers without being had
Three systematic flatteries. Baseline choice: “35× lower than Hopper” compares against a two-generation-old platform at its worst workload fit (long-context reasoning), not against the B200 fleet you might otherwise buy. Workload fit: multipliers peak on high-batch, FP4-quantised, long-reasoning serving that exploits the 72-GPU NVLink domain; short-context chat on a 70B model shows far smaller deltas — and may not need rack scale at all, per rack-scale vs scale-out. Utilisation assumption: headline cents-per-million-token figures presume the rack is saturated. None of this makes the platform bad — it makes your own benchmark, on your model mix, the only number worth signing against.
Break-even against APIs and cloud rental
Directional 2026 arithmetic: frontier-model APIs bill on the order of $1–15 per million output tokens depending on model class; open-weight serving via APIs runs $0.1–1. An owned NVL72-class rack serving open-weight models at 50% utilisation typically lands in the low cents per million tokens on optimised stacks — one to two orders of magnitude under API list — before counting the sovereignty and DPDP benefits of in-country serving. The catches: capex is committed before demand materialises; engineering an optimised serving stack (TensorRT-LLM/vLLM, disaggregation, quantisation) is real work; and below sustained mid-double-digit utilisation, cloud rental at $2–3/GPU-hour or IndiaAI-subsidised capacity remains the cheaper bridge. The sequencing in our inference sizing checklist — prove demand rented, then buy the baseline — applies with extra force at rack scale.
Sensitivity: what moves the token cost
| Lever | Swing | Effect on ₹/M tokens | Owner |
|---|---|---|---|
| Utilisation 30% → 70% | 2.3× more tokens | −57% | Demand planning, consolidation |
| Precision FP8 → FP4 (where quality holds) | ~1.5–2× throughput | −33–50% | Serving engineering |
| PUE 1.6 → 1.25 | −22% power bill | −5–8% | Facility choice |
| Amortisation 36 → 60 months | −40% monthly capex line | −20–30% | Finance policy + mid-life role plan |
| Batch/disaggregation tuning | 1.5–3× throughput | −33–66% | Serving engineering |
Frequently asked questions
What token cost should an owned rack achieve?
Serving open-weight models on an optimised stack at 50%+ utilisation, expect low single-digit cents per million tokens for mid-size models and tens of cents for the largest — but derive it from your own benchmark, tariff and utilisation, not from marketing slides.
Are the 15×-ROI and 35×-cheaper claims false?
Not false — conditional. They model saturated racks, FP4-optimised reasoning workloads and (for ROI) API-resale revenue. Enterprises avoiding API spend at partial utilisation see large but smaller advantages.
Why do power metrics matter more than price per GPU?
Because liquid-cooled kilowatts are the scarce resource in 2026 facilities, including India’s. Tokens per megawatt determines what a fixed facility ceiling can earn — and newer generations win primarily on that axis.
When does API spend justify buying a rack?
A steady seven-figure-rupee monthly API bill for open-weight-servable workloads is the usual tripwire — at that run-rate, an owned or hosted rack with a competent serving team typically pays back within its first two years. Verify with the five-line model at conservative utilisation.
Does this analysis apply to fine-tuning and training?
Partially — training economics are about time-to-result, not tokens served, and favour rental for bursts. The token model governs the serving estate, which is where rack-scale systems spend most of their life.
Ready to deploy?
Talk to an RDP architect about power, cooling and lead time.