TheVortiq
Inteligencia Artificial

NVIDIA Vera Rubin: Record Performance and Reduced Token Cost

The new platform promises 10x more tokens per megawatt than Blackwell, according to CoreWeave tests

July 24, 2026 · 4 min read

blue UTP cord

TL;DR: NVIDIA presents Vera Rubin, an integrated AI platform achieving 10x more tokens per megawatt than Blackwell, according to CoreWeave benchmarks. This drastically reduces cost per token for partners.

What happened?

NVIDIA has announced the Vera Rubin platform, a complete AI system from chip to power grid, designed to maximize performance per watt and minimize cost per token. Production of Vera Rubin NVL72 is already underway in over 350 factories across 30 countries, with partners including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. According to NVIDIA's official blog, the platform integrates seven chips and five trays — Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX — all designed as a single system, not as assembled components from standard products. This extreme vertical integration allows optimization of performance per watt and cost per token, two critical metrics in the era of large-scale AI.

Why is it important?

The key data comes from a CoreWeave benchmark on DeepSeek-R1: Vera Rubin delivers 10x more tokens per megawatt than Grace Blackwell NVL72. This is critical because energy consumption is the main bottleneck for expanding AI data centers. Additionally, the platform incorporates the sixth generation of NVLink, which doubles performance on complex workloads, reduces latency by 3x, and multiplies packet rates by 10x compared to standard Ethernet. For out-scale, Spectrum-X combines 102.4T switches and 1.6T SuperNICs, offering 1.6x more RDMA bandwidth. The integration of photonics with co-packaged optics reduces power consumption by 5x and increases reliability (MTBI) by 10x compared to pluggable transceivers, lowering maintenance costs. This leap in energy efficiency is vital: according to the International Energy Agency, AI data centers could consume up to 100 TWh by 2026, equivalent to the consumption of a country like Sweden.

Market implications

Vera Rubin redefines efficiency in large-scale AI. For companies like CoreWeave, Microsoft, and Tesla, which are already adopting Spectrum-6 switches, this means being able to scale their AI factories without skyrocketing energy costs. The reduction in cost per token will allow cloud providers to offer more competitive prices, accelerating AI adoption in sectors like healthcare, finance, and automotive. In terms of competition, NVIDIA strengthens its leadership against alternatives like AMD or custom solutions from Google and Amazon. However, the AI accelerator market is valued at over $100 billion by 2027, according to Gartner, so the pressure on NVIDIA to maintain its edge is enormous. The Vera Rubin platform could force competitors like AMD to accelerate their own integrated designs or risk losing market share.

What readers should know

Vera Rubin is not just a chip; it's a complete platform that includes the Vera CPU, with Olympus cores offering 2x single-thread performance, 3x inter-core bandwidth, and 40% lower memory latency than competing designs. This makes it ideal for AI agent workloads, which require low latency and high efficiency. For partners, the lower cost per token translates into more competitive prices for end users of cloud AI services. Additionally, photonics integration reduces cabling complexity and improves reliability, which is crucial for data centers operating 24/7. The sixth-generation NVLink also allows scaling to thousands of GPUs with low latency, facilitating the training of massive models like large language models.

"The Vera Rubin platform is built from chip to network to deliver the highest performance per watt and the lowest cost per token," NVIDIA states in its official blog.

Historical context

Compared to Blackwell, Vera Rubin represents a generational leap in efficiency. While Blackwell focused on raw performance, Vera Rubin optimizes performance per watt, something crucial at a time when electrical capacity is the scarcest resource for data centers. The adoption of photonics and sixth-generation NVLink shows NVIDIA's bet on vertical integration and extreme co-design. Historically, NVIDIA has gone from being a GPU manufacturer for gaming to dominating the AI market with architectures like Volta (2017), Turing (2018), Ampere (2020), Hopper (2022), and Blackwell (2024). Each generation has roughly doubled performance per watt, but Vera Rubin accelerates this trend with a 10x factor in token efficiency per megawatt. This focus on energy efficiency recalls the shift that the Arm architecture brought to mobile devices, where performance per watt became the key metric. If NVIDIA maintains this pace, it could set a standard that forces the entire industry to rethink AI data center design.

Keep reading