The Information Machine
Following·since 21 Aug 2026·Day 6·3 sources·updated 27 Aug 2026

NVIDIA Vera Rubin and NVLink Fusion platform

The gist

NVIDIA Vera Rubin Ships to Azure; NVLink Fusion Adds NVHBM

Vera Rubin arriving at Azure marks the start of production deployment for NVIDIA's latest GPU generation at a major cloud provider. NVLink Fusion and NVHBM give hyperscalers building custom silicon a path to integrate with NVIDIA's networking and rack-scale infrastructure, with Amazon's Trainium4 as the first announced NVHBM customer.

The full picture

Microsoft Azure has taken delivery of the first production NVIDIA Vera Rubin GPU systems. NVIDIA published early benchmark data showing Vera Rubin NVL72 achieving up to 30x higher inference throughput per megawatt and 35x lower cost per million tokens compared to GB300 NVL72, measured using the SemiAnalysis AgentX benchmark on agentic coding workloads; those results are pending SemiAnalysis review and do not yet reflect Vera CPU performance for tool calling. NVLink Fusion connects custom third-party XPUs to NVIDIA's sixth-generation NVLink scale-up networking, and Scale-In was announced as a fifth AI networking pillar powered by BlueField-4 DPUs. NVIDIA also expanded NVLink Fusion with NVHBM, a custom memory technology that integrates NVIDIA's memory controller into the HBM base die, with Amazon's Annapurna Labs named as the first partner for use in Trainium4 chips.

How it developed
27 August 2026

NVIDIA confirms Vera CPU shipping at scale; AWS receives first Vera CPU server and Vera Rubin GPU; Phoronix finds Vera 10% ahead of AMD EPYC 9575F on geometric mean across permitted workloads

NVIDIA on August 27 announced NVHBM as part of NVLink Fusion, a custom memory architecture that embeds the memory controller in the HBM base die rather than on the XPU die, freeing silicon area for compute. NVHBM delivers up to 30% greater bandwidth and 15% lower power than standard HBM4E while freeing up to 25% more XPU die area. Amazon's Annapurna Labs became the first named partner, committing to support NVHBM in its Trainium4 chips.

First citedNVIDIA Blog
26 August 2026

NVIDIA expands NVLink Fusion with NVHBM; Amazon's Annapurna Labs named as first partner for Trainium4

The Groq 3 LPX processor uses deterministic compiler scheduling and 128GB of SRAM across the rack to keep decoding delays from compounding across sequential steps in agentic workflows, with preplanned chip-to-chip transfers reducing coordination overhead for small batches. Reporting dated August 26 placed the processor alongside the Vera Rubin NVL72 systems Azure received in August 2026 and NVIDIA's NVLink Fusion platform, citing the Wall Street Journal on agentic AI posing two distinct computing challenges: processing large contexts efficiently and generating tokens with low latency.

25 August 2026

SpaceXAI announced plans to deploy NVIDIA's Vera CPU for Grok and fly Vera hardware on the Starmind satellite

Microsoft Azure took delivery of the first production NVIDIA Vera Rubin GPU systems in August 2026, with Microsoft crediting both NVIDIA and its own hardware and datacenter teams, while NVIDIA published benchmark claims that Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads, figures NVIDIA states are pending SemiAnalysis review. NVIDIA also announced NVLink Fusion, connecting custom XPUs and CPUs to sixth-generation NVLink with Intel, QCT, Mediatek, and Annapurna Labs as named participants, and Scale-In as the fifth pillar of its AI networking lineup.

24 August 2026

NVIDIA announces NVLink Fusion and Scale-In details including BlueField-4 and Spectrum-X integration

22 August 2026

NVIDIA acquires minority stake in Cloverleaf, a data center land and power developer

21 August 2026

Microsoft Azure receives first production NVIDIA Vera Rubin GPU systems

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free