Nvidia opens its networks to custom processors: a unified architecture to link third-party chips with AI factories
Listen to this article
Read by Anchor
Companies that build custom AI accelerators understand that chip design alone is only part of the complex equation of constructing modern compute centers, as AI factories require an integrated infrastructure that includes vertical and horizontal scaling networks, rack-level server architectures, liquid cooling systems, and operating software stacks. In this context, Nvidia unveiled the NVLink Fusion technology, an architectural initiative that enables third-party custom compute processors (XPUs) to connect to the infrastructure and networking system that the company develops, reducing operational risk and accelerating time-to-market for hybrid compute installations.
The economics of AI factories depend on strict operational metrics tied to tokens generated per second per watt, the cost per token, and device readiness and actual utilization rates. As workloads grow more complex, such as recursive reasoning models, trillion-parameter models, and mixture-of-experts (MoE) architectures, vertical scaling networks become the primary bottleneck, because a network that cannot keep pace with processors leads to a sharp drop in utilization efficiency and a rise in cost.The new technology provides connectivity for custom processors within the sixth generation of NVLink, reducing inter-processor data latency by a factor of three and increasing packet rates by tenfold compared with conventional Ethernet solutions.
The system integrates via the NVLink-C2C interface to link custom processors with CPUs, achieving power-efficiency six times higher than PCIe interfaces. This pathway allows chip developers such as Intel, MediaTek and Amazon’s Annapurna Labs to leverage the reference designs for MGX server racks and the architecture used in Vera Rubin NVL72 systems, which employ 100 % liquid cooling without the need for fans or complex cabling inside compute drawers, while permitting drawer maintenance and replacement without interrupting server operation.
The significance of this architecture extends beyond hardware to the software and runtime layer, as the technology integrates NCCL libraries for workload distribution, Dynamo and NIXL tools for resource partitioning and allocation, and the Mission Control platform for cluster management and remote measurement. Omniverse DSX designs also provide reference digital twins of gigawatt-scale compute factories, allowing engineers to test power and cooling flows and plan infrastructure before actual construction begins.
This approach changes the operational calculations for large infrastructure projects and data centers in the Gulf and Egypt,Investors in massive compute facilities face the dilemma of committing early to a single type of chip before data centers are completed and power and cooling sources are secured. The unified architecture allows facility managers and purchasers to pre-install 800-volt DC power distribution networks and to design liquid-cooling systems to standard dimensions that support both conventional GPUs and sovereign or custom processors, freeing operators from strict dependency constraints and enabling full flexibility in reallocating compute capacity as supply chains evolve and model requirements change.
This step places the high-performance computing industry at a stage where engineering value differentiates: instead of expending resources to reinvent interconnects and rack systems from scratch, chip innovators can focus on designing the specialized processor, while standard networks provide the speed and stability needed to run systems with continuous industrial efficiency.