Skip to main content

NVIDIA Vera Rubin Enters Full Production: The Agentic AI Factory Era Begins

· 5 min read
Industry Research Team

On June 1, 2026, NVIDIA founder and CEO Jensen Huang officially announced at COMPUTEX 2026 (Taipei) that: the Vera Rubin platform has entered full production. This marks a fundamental paradigm shift for AI hardware from "discrete accelerators" to "integrated AI factories."

Key Highlights

  • Rubin GPU: Next-gen AI compute chip, FP4 compute is 3.6× that of Blackwell
  • Vera CPU: 88 custom Arm cores (176 threads), replacing the Grace CPU
  • NVLink 6: GPU-to-GPU interconnect bandwidth reaches 260 TB/s (double Blackwell)
  • CX8 SuperNIC: 800Gb/s network, ConnectX-9 link reaching 28.8 TB/s
  • HBM4 memory: 288GB per chip, 13 TB/s bandwidth
  • Agentic throughput: 10× over Grace Blackwell

Complete Vera Rubin Platform Specs

Vera Rubin is not a single GPU but a complete AI factory platform comprising 7 chips:

ChipTypePurpose
Rubin GPUMain AI compute chipTraining + inference
Rubin Ultra GPUFlagship versionUltra-scale inference
Vera CPUCPU paired with RubinHost CPU + data preprocessing
NVLink 6Interconnect chipHigh-speed GPU interconnect (260 TB/s)
CX8 SuperNICNIC800Gb/s network
XDR 800G switchDatacenter networkCross-rack communication
Rubin Platform PODWhole cabinetPre-configured AI factory (144 GPUs)

Rubin GPU Detailed Specs (estimated)

ParameterRubin GPURubin UltraBlackwell (B200)
ArchitectureRubinRubin UltraBlackwell
ProcessTSMC 3nm (est.)TSMC 3nmTSMC 4NP
Memory288GB HBM4288GB HBM4E (est.)192GB HBM3e
Memory bandwidth13 TB/s13+ TB/s8 TB/s
FP4 compute~3,600 TFLOPS (est.)~5,000 TFLOPS (est.)2,250 TFLOPS
TDP1,000W (est.)1,200W (est.)700-1000W
InterconnectNVLink 6 (260 TB/s)NVLink 6NVLink 5 (1800 GB/s)
Mass production2026 Q3H2 20272024 Q4

📌 Note: Rubin's exact specs are not fully public yet; some values above are estimates.

Vera CPU: The New Host CPU Replacing Grace

Vera CPU is NVIDIA's self-designed Arm-architecture CPU, replacing the previous Grace CPU:

ParameterVera CPUGrace CPU
Cores88 cores (176 threads)72 cores (144 threads)
ArchitectureCustom Armv9 (est.)Arm Neoverse V2
InterfaceNVLink 5.0 (1.8 TB/s)NVLink 4.0 (900 GB/s)
TDP~500W (est.)350-500W
PurposeAI factory Host CPUHPC / AI Host

Key upgrade: Vera's co-design with the Rubin GPU achieves end-to-end optimization in compute, data loading, and preprocessing, comparable to Google TPU 8t's Arm Axion integration.

Performance vs Blackwell

NVIDIA officially claims that under the same POD configuration (144 GPU chips):

MetricGrace Blackwell (GB200 NVL72)Vera Rubin NVL144Improvement
FP4 compute1.1 PFLOPS3.6 PFLOPS3.3×
Memory capacity288GB×72 = 20.7TB288GB×144 = 41.4TB
Memory bandwidth8 TB/s×7213 TB/s×144~3.3×
NVLink bandwidth1800 GB/s×72260 TB/s (full POD)~2×
Agentic throughputBaseline10×10×
Performance per wattBaseline25× (vs CPU alone)25×

💡 Why "10× agentic throughput"? Agentic AI workloads differ from training/inference: one prompt may trigger multiple stages including reasoning, retrieval, tool calls, and response generation, involving thousands of steps. The Rubin platform is optimized for this long-chain, high-concurrency workload.

MGX Third-Gen Rack-Scale System

Vera Rubin adopts the MGX third-gen open rack-scale system design:

  • Five-rack synergy: Vera Rubin NVL72 system + Vera CPU + Groq 3 LPX + Vera BlueField-4 STX storage + Spectrum-6 SPX Ethernet
  • Global supply chain: 30 countries, 350+ factories, hundreds of partners (Dell, HPE, Lenovo, Supermicro, Asus, Foxconn, etc.)
  • Spectrum-X Ethernet silicon photonics: World's first switch based on CPO (co-packaged optics) supporting 200Gb/s SerDes, now in mass production

Mass Production Timeline

TimeEvent
Jan 2026CES 2026 first unveils Rubin platform
June 1, 2026COMPUTEX 2026 announces full production
Fall 2026Vera Rubin officially starts mass production and shipment
H2 2027Rubin Ultra launch (HBM4E upgrade)
2028Feynman architecture (next gen)

AI Factory: From Selling Chips to Selling "Smart Production Lines"

Huang said something at the launch that shook the industry:

"Rubin's Agentic AI throughput is 10× that of Blackwell. Rubin is a complete AI factory platform."

This marks a fundamental shift in NVIDIA's business model:

  • Past: Sold GPUs (H100/B200), customers built systems themselves
  • Now: Sells "complete AI factory solutions" (Vera Rubin POD), including GPU, CPU, network, storage, software stack
  • Future: Becomes the "TSMC" of global AI infrastructure (providing smart production capacity)

vs Competitors

VendorProductPositioningAdvantageDisadvantage
NVIDIAVera RubinComplete AI factory solutionMost complete ecosystem, most mature softwareExpensive, extremely high power
AMDMI455X (MI400 series)Training competitorPrice/performance, open ecosystemSoftware ecosystem gap
GoogleTPU 8i/8tCloud training/inferenceDeep Gemini integrationGoogle Cloud only
HuaweiAscend 910C/950Domestic substitutionChina localization, AscendMind frameworkAffected by export controls

Industry Impact

  1. AI labs: Frontier model training time shrinks from "months" to "weeks"
  2. Cloud providers: Must decide whether to procure Vera Rubin POD (conflicts with self-developed chip strategy)
  3. Hyperscale datacenters: AI factory becomes a new competitive dimension (whoever has the strongest compute can train the strongest model)
  4. Domestic chips: Ascend 910C/950, Cambricon MLU590, etc. must catch up to Blackwell in 2026-2027, or the gap will widen to the Rubin era

References


This article is compiled from NVIDIA official announcements and public materials. Some specs are estimates, subject to final official release.