Skip to content
Make in India OEM · INR-transparent · Pan-India onsite SLATalk to sales: +91 720 794 8743Sign in

Blackwell Ultra to Vera Rubin to Feynman: The 2026–2028 AI Training Cluster Roadmap

Updated 15 Jul 2026 · 6 min read

Overview

NVIDIA is now shipping on a one-architecture-per-year cadence, and that single fact should shape how you plan a training cluster between now and 2028. Blackwell Ultra (GB300) is the platform you can buy and run today; Vera Rubin entered production in mid-2026 with partner systems arriving in the second half of the year; Rubin Ultra follows in H2 2027 and Feynman in 2028. The practical question for most teams is no longer “which GPU is fastest” but which generation your facility can actually power and cool, and whether waiting one cycle costs more than it saves.

Blackwell Ultra to Vera Rubin to Feynman: The 2026–2028 AI Training Cluster Roadmap
What you’ll learn: how the Blackwell → Rubin → Rubin Ultra → Feynman cadence is sequenced, what Vera Rubin’s HBM4 memory changes for large training, why rack power (370 kW → 1 MW+) and not silicon is the binding constraint, a three-test frame for deciding buy-now vs wait, and what it means for India buyers under IndiaAI and DPDP.

Key takeaways

  • Annual cadence is the planning assumption: Blackwell → Rubin (H2 2026) → Rubin Ultra (H2 2027) → Feynman (2028).
  • Blackwell Ultra (GB300) is the buy-today platform — roughly 1.5× the AI performance of GB200 NVL72.
  • Vera Rubin (VR200) — reported at ~288 GB HBM4 per GPU and ~50 PFLOPS FP4, in production since mid-2026.
  • Facility limits, not silicon, are the real constraint: ~370 kW racks in 2026, Kyber ~600 kW in 2027, 1 MW+ by 2027–2028.
  • Waiting is not free: a generation you cannot power, cool or get allocation for is not a plan.

The cadence: Blackwell, Rubin, Rubin Ultra, Feynman

NVIDIA has moved to a yearly architecture rhythm, and the publicly discussed sequence runs Blackwell → Rubin (second half of 2026) → Rubin Ultra (second half of 2027) → Feynman (2028). Under a two-year cadence you could reasonably wait for the “next one”; under an annual cadence there is always a next one roughly twelve months out, so waiting indefinitely means never deploying. The useful planning horizon becomes: buy the generation that fits the facility you have or are building, and design the site so the following generation can land without a rebuild.

What each generation changes

  • Blackwell Ultra (GB300 NVL72) — the shipping rack-scale platform, reported at roughly 1.5× the AI performance of GB200 NVL72, with strong MLPerf Inference v5.1 results. Mature systems, real benchmark data, settled supply chain: the lowest-deployment-risk choice for capacity needed this financial year.
  • Vera Rubin (VR200) — the successor, in production since mid-2026 with partner availability in H2 2026. Reported at ~288 GB of HBM4 and ~50 PFLOPS FP4. The headline is memory as much as compute: more HBM per GPU means larger model shards and longer context per device, reducing the parallelism gymnastics that dominate large-training engineering.
  • Rubin Ultra (Kyber NVL576) — expected H2 2027, associated with roughly 600 kW per rack. This is a facility-led decision long before it is a silicon one.
  • Feynman — 2028, the next cadence step. Plan the site, not the SKU.

Roadmap at a glance

Generation Timing What it signals for buyers
Blackwell Ultra (GB300 NVL72) Shipping (mid-2026) ~1.5× GB200 NVL72; mature, lowest deployment risk
Vera Rubin (VR200) Production mid-2026; partner systems H2 2026 ~288 GB HBM4, ~50 PFLOPS FP4; bigger memory per GPU
Rubin Ultra (Kyber NVL576) H2 2027 Rack design targeting ~600 kW — facility-led decision
Feynman 2028 Next cadence step; plan the site, not the SKU

The real constraint is power and cooling

Each step up the roadmap raises rack power faster than most data centers can absorb. Next-generation AI racks are projected around 370 kW in 2026, Rubin Ultra’s Kyber design is associated with roughly 600 kW per rack in H2 2027, and single racks are widely expected to pass 1 MW in the 2027–2028 window. Two consequences follow. First, direct-to-chip liquid cooling is no longer optional — the market has moved past the experimental phase and is consolidating around DLC as the default for AI-centric deployments, valued near $3.7 billion in 2026 and projected to roughly $18.1 billion by 2036. Second, power distribution is shifting to 800 VDC: Vertiv, Schneider Electric, Eaton and Delta have signalled commercial products in H2 2026, timed to the Kyber rack window. Higher voltage means lower current for the same power — thinner conductors, less copper in the rack, fewer conversion stages. Adoption stays gradual, with industry estimates putting 800 VDC at only 15–25% of facilities by 2030. If your site is AC-distributed and air-cooled today, that gap — not the GPU roadmap — is your binding constraint.

Buy now or wait: a three-test frame

Use three tests rather than a launch calendar. One: can your site power and cool the generation you want? A GB300-class deployment you can energise beats a Rubin-class order your facility cannot host. Two: what is the cost of idle time? Under an annual cadence, a year spent waiting is a year of training not done. Three: can you get allocation? Top-of-roadmap platforms are frequently constrained, with supply prioritised for strategic partners — availability, not preference, often decides. Where all three pass for the newer generation, wait. Where any fails, deploy the mature platform now and design the site so the next generation lands without a rebuild.

What this means for India buyers

India’s demand signal is strong and specific. The IndiaAI Mission has empanelled more than 38,000 GPUs, targeting roughly 100,000 by the end of 2026, backed by about $1.25 billion (₹10,372 crore) with subsidised access near $1 per GPU-hour. That makes rented capacity attractive for experimentation, but it does not resolve data residency. Under the DPDP Act, personal-data handling carries penalties up to ₹250 crore per violation; the Act stops short of blanket localisation but pushes regulated workloads toward in-country, controlled infrastructure. The practical pattern for BFSI, healthcare and public-sector teams is hybrid: burst experimentation on subsidised or cloud capacity, and keep sensitive training and fine-tuning on owned, in-country clusters where residency and audit are provable.

Frequently asked questions

Should I wait for Vera Rubin instead of buying Blackwell Ultra?

Only if your facility can power and cool it and you can secure allocation. Under an annual cadence there is always a newer generation ~12 months out, so waiting on principle means never deploying. If your site is ready and allocation is available, Rubin’s larger HBM4 memory is a genuine advantage; otherwise deploy GB300 now.

What comes after Vera Rubin?

Rubin Ultra is expected in the second half of 2027, associated with the Kyber NVL576 rack design and roughly 600 kW per rack, followed by Feynman in 2028.

How much power will an AI rack need by 2028?

Projections put next-generation racks near 370 kW in 2026, around 600 kW for Rubin Ultra Kyber in 2027, and above 1 MW in the 2027–2028 window. Direct-to-chip liquid cooling and 800 VDC distribution are the enabling technologies.

Do I need 800 VDC and liquid cooling now?

Liquid cooling — yes, in practice, for any current rack-scale AI deployment. 800 VDC is arriving commercially in H2 2026 but adoption is gradual (estimated 15–25% of facilities by 2030), so treat it as a design target for your next build rather than a prerequisite today.

Is renting IndiaAI GPU capacity enough for regulated workloads?

It suits experimentation and burst training at subsidised rates. For workloads touching personal data under DPDP, most regulated teams keep training and fine-tuning on owned in-country infrastructure where residency, access control and auditability can be demonstrated.

Ready to deploy?

Talk to an RDP architect about power, cooling and lead time.

Request a Quote