Nvidia integrates Groq 3 accelerators into the Vera Rubin platform with an integrated inference architecture to reduce generation costs in agent systems
Listen to this article
Read by Anchor
At the Hot Chips conference in California, Nvidia announced that its new Groq 3 LPX system has entered full commercial production, integrated with the Vera Rubin NVL72 supercomputing platform, in a shift redesigning data centres from model training clusters into what it terms token factories tailored for autonomous systems and intelligent agents.
Independent benchmarks conducted by Artificial Analysis on the open-source Gemma 4 31-billion-parameter model showed the new platform generating 3,400 tokens per second across long contexts of up to 100,000 tokens, four times the speed of the nearest alternative on the market.The engineering rationale rests on a strict division of workloads, where Rubin GPUs process massive contexts while LPX units accelerate the decoding phase and sequential token generation.
The Groq 3 LPX architecture relies on 256 LP30 accelerators linked by direct die-to-die interconnects within a single rack, operating as a single giant processor dedicated to low-latency, deterministic inference. Cloud provider Nebius has begun deploying these accelerators in its infrastructure to allow developers to build coding agents and real-time interactive experiences, while SpaceX AI announced it is adopting Vera CPUs to orchestrate agent tasks, tool use, and code execution from terrestrial data centres to orbital satellites.
To address networking bottlenecks at hyperscale, Nvidia introduced its Spectrum-X Multi-Plane technology, already deployed by CoreWeave. The network architecture splits server connections into independent planes, enabling scaling up to 512,000 GPUs across a flat two-tier network without requiring a third tier that increases latency and cost. The network retains roughly 90% of its total capacity if a plane fails, with hardware-level recovery that is 11 times faster than software-based alternatives.
The company also unveiled the Scale-In platform, powered by BlueField-4 DPUs and DOCA software to accelerate security services and shared infrastructure management, alongside NVLink Fusion, which allows companies to integrate custom silicon and processors into the unified system using sixth-generation NVLink interconnects and open MGX standards.
This transition from raw training throughput to decoding economics is reshaping cloud infrastructure and regional data centre planning across the Gulf and Egypt.The shift to running intelligent agents means that the cost of token consumption across long contexts and the speed of response have become the true measure of commercial viability. Lowering generation latency while supporting massive contexts enables local enterprises and institutions to run complex automation workflows for customer service and financial analysis without exhausting compute budgets on processors unsuited for iterative decoding.