Infrastructure

NVIDIA Vera Rubin in Production: AI Factories, 10x Per-Watt Inference

NVIDIA used its GTC Taipei and Computex stage to declare its next-generation Vera Rubin platform in full production and to push a single idea: the AI race is moving from chatbot windows into rack-scale "AI factories." The pitch targets agentic AI, reasoning, and long-context workloads — the workloads that run for minutes, not milliseconds. For anyone building, buying, or budgeting AI, this is the layer that decides cost per token. It is also a reminder that frontier AI now depends as much on power, cooling, and supply chains as on model weights.

· Jun 15, 2026 · updated Jun 16, 2026
NVIDIA Vera Rubin in Production: AI Factories, 10x Per-Watt Inference
Table of contents
  1. What changed
  2. Why it matters and for whom
  3. What to watch and what's next
  4. Bottom line

What changed

NVIDIA announced that its Vera Rubin platform is in full production. The flagship building block is the Vera Rubin NVL72, a rack that pairs 36 Vera CPUs with 72 Rubin GPUs. NVIDIA frames the design around five integrated racks, adding the Vera CPU, Groq 3 LPX accelerators, Spectrum-X Ethernet photonics for networking, and BlueField-4 STX storage.

The headline numbers are about efficiency, not raw speed. NVIDIA claims 10x higher inference performance per watt and 10x lower cost per token versus the prior generation. Paired with Groq 3 LPX, it cites up to 35x higher throughput per watt for trillion-parameter models, and up to 10 petaflops of AI performance.

The engineering story is industrial. NVIDIA describes a cable-free, hose-free, fanless system using 100% liquid cooling with a 45-degrees-Celsius warm-water inlet. It says compute-tray assembly drops from two hours to five minutes, and that power shelves carry 6x more onboard energy storage. Each system, NVIDIA says, contains almost 2 million parts.

Scale is the other point. NVIDIA says the Vera Rubin supply chain is twice as large as Grace Blackwell, spanning 350-plus factories across 30 countries, with 150 ecosystem partners in Taiwan alone.

Why it matters and for whom

The explicit target is agentic AI, reasoning, and long-context workloads. That choice of words matters. A simple chatbot reply is cheap. An agent that plans, calls tools, and works through a long task burns far more compute per request. Reasoning models think in many internal steps. Long-context jobs hold huge prompts in memory. All three push the bottleneck from one chip to the whole rack — and then across the data center.

This is why NVIDIA keeps repeating the phrase AI factory. The unit of value is shifting from a single GPU to an integrated rack of compute, networking, storage, and cooling, measured by output per watt and cost per token. For hyperscalers and labs, that reframes the buying decision around total cost of inference, not benchmark peaks.

The Stanford AI Index 2026 gives the backdrop: inference is now the dominant cost as adoption scales, and efficiency gains compound directly into margins. If you serve millions of users, a 10x improvement in performance per watt is the difference between a viable product and a money pit.

The efficiency framing is also a constraint disguised as a feature. Liquid cooling and dense energy storage exist because power and heat are the real ceilings. Data centers are increasingly limited by grid access, not chip availability. NVIDIA's pitch quietly concedes that the frontier is now an energy and supply-chain problem.

The assembly numbers point the same way. Cutting compute-tray assembly from two hours to five minutes is not a vanity stat; it is throughput. When a single system holds almost 2 million parts and demand outstrips supply, manufacturing speed becomes a competitive weapon. The hose-free, cable-free design exists to make racks faster to build and easier to service at scale — an admission that the bottleneck has moved from designing chips to shipping finished, working racks. In an AI factory, the production line matters as much as the silicon on it.

There is a lock-in dimension too. By integrating compute, the Vera CPU, Groq 3 LPX accelerators, Spectrum-X networking, and BlueField-4 storage into one platform, NVIDIA sells a system, not a part. That makes the per-watt and per-token gains real, but it also makes it harder for buyers to mix vendors. The same integration that delivers efficiency deepens dependence on a single supplier for the whole stack.

What to watch and what's next

Production claims are a starting line, not a finish. The questions are how fast racks reach customers, whether the cited efficiency holds on real agentic workloads rather than ideal benchmarks, and whether power and cooling capacity can keep pace with demand.

Watch the supply chain. A claim of twice the scale of Grace Blackwell implies enormous coordination across factories and countries; that is exactly where ramps slip. Watch competitors and custom silicon from cloud providers, who all want to escape paying the cost per token NVIDIA sets. And watch energy: warm-water cooling and onboard storage hint that the next constraint on AI growth is the grid.

For buyers, the practical move is to model cost per token and per-watt throughput for your own workloads before committing. For everyone else, Vera Rubin is the clearest sign yet that the biggest AI decisions are being made in racks, not in chat windows.

Bottom line

  • NVIDIA says Vera Rubin is in full production, built as a rack-scale AI factory for agentic AI, reasoning, and long-context workloads.
  • The pitch is efficiency: 10x inference performance per watt and 10x lower cost per token, with 100% liquid cooling and a supply chain twice the size of Grace Blackwell.
  • The competitive frontier is shifting from models to infrastructure — power, cooling, supply chains, and the price of a token.

The next phase of the AI race will likely be won or lost on whether these factories can be powered, cooled, and shipped fast enough to meet demand.

More on NVIDIA