Unified rack architecture opens AI factory doors as Nvidia integrates De Matrix chips to accelerate inference
Listen to this article
Read by Anchor
The race to accelerate AI computing is moving toward a new stage in which the chip is no longer separate from the hosting infrastructure, with De Matrix, a specialist in inference chips, announcing adoption of Nvidia's NVLink Fusion platform to integrate its next-generation Raptor processors directly into the high-performance computing platform and AI factory servers. This alliance enables custom silicon developers to leap over the industry's toughest hurdle: moving the processor from theoretical design to actual operation inside liquid-cooled giant server racks.
The step stems from an explicit operational insight expressed by Mr. Seth, co-founder and CEO of De Matrix, who noted that demand for inference capability is experiencing major jumps while capital, time and energy resources remain limited and subject to strict constraints. Seth affirmed that linking with the liquid-cooled MGX rack architecture gives customers a faster, lower-risk path to deploying inference processors with ultra-low latency and scaling them without having to devise cooling or power-distribution solutions from scratch.
Designing the innovative chip represents only half the challenge, while the heavier half lies in operationally integrating it within an established supply, cooling and networking system.The NVLink Fusion platform relies on opening Nvidia’s interconnect architecture to custom processors and central processors based on ARM, x86 and RISC-V architectures. According to the platform’s technical data, the sixth generation of NVLink delivers inter-processor latency three times lower than commercial Ethernet solutions on the market, with packet throughput ten times higher and a total inter-processor bandwidth of up to 3 TB/s per processor.
De Matrix plans to connect its processors in a ultra-fast, unified, scalable range that allows its specialized racks to operate alongside Nvidia GPU systems such as the Vera Rubin NVL 72 for distributed inference tasks. The plans include integrating central Vera processors and Connect X-9 SuperNIC network cards, as well as Bluefield-4 data-processing units and Spectrum-X Ethernet solutions, joining a broad partner list that includes AWS, Intel, Samsung, Arm, Marvell and Cadence.
This shift directly impacts data-center building plans in the region, especially in the Gulf states and Egypt, where organisations race to construct high-performance computing facilities with liquid cooling that require massive capital investment. The practical value lies in infrastructure flexibility: adopting a unified, open rack architecture for multiple chips means a local operator will not have to allocate separate halls for each processor type or redesign cooling and power-distribution systems when testing specialized inference chips that lower the cost per token. Engineering teams and national cloud operators can insert specialized inference accelerators alongside existing graphics hardware without risking supply-chain disruption or incurring duplicate capital expenses to retrofit infrastructure for each new model entering the market.
Standardising the interconnect language and racks turns custom inference processors from isolated islands into additional gears rotating within the prevailing computing ecosystem.We see in this a resolution of the traditional trade-off between owning ultra-efficient specialized silicon and the complex deployment requirements of large data centers, as these platforms give chip makers an immediate bridge to market while ensuring operating facilities can preserve their existing infrastructure investments without fundamental modification.