# What happened AI-infrastructure startup d-Matrix will license Nvidia's NVLink Fusion interconnect and adopt MGX rack-scale reference designs for its next-generation Raptor accelerators. The move lets d-Matrix build systems that plug into the same rack and NVSwitch fabrics used by Nvidia, and to deliver racks with up to 144 accelerators tied together by a single all-to-all NVLink fabric.
# Why it matters Scaling an accelerator architecture across large clusters requires a high-speed, low-latency fabric. By using NVLink Fusion and MGX rack designs, d-Matrix avoids designing its own rack interconnects and leverages an established ecosystem of NVSwitch appliances, Vera CPUs, BlueField and ConnectX NICs, and SpectrumX Ethernet. That reduces integration work for customers who already deploy Nvidia-compatible racks.
# Technical snapshot Raptor accelerators use bonded compute logic on top of 3D-stacked DRAM to push memory bandwidth close to SRAM-like levels. At Hot Chips the company showed a full Raptor card with 32 GB of 3D-stacked DRAM capable of about 100 TB/s of memory bandwidth. d-Matrix's NVL144 rack concept uses smaller XPUs—roughly half the size of the Hot Chips Raptor card—with about 16 GB of 3D-DRAM and ~50 TB/s of per-XPU bandwidth.
With 144 XPUs per rack, the design targets approximately 2.3 TB of total memory capacity and a peak aggregate memory bandwidth near 7.2 petabytes per second. That memory profile is pitched at inference for very large models: d-Matrix estimates the NVL144 configuration can serve models exceeding four trillion parameters at 4-bit precision.
# How this fits inference workflows
# Business and ecosystem effects
# What to watch next
- d-Matrix timeline for shipping NVL144 systems and availability of the full Raptor cards.
- Real-world throughput and latency metrics for heterogeneous GPU+Raptor inference pipelines.
- Any commercial or investment terms disclosed that clarify Nvidia's incentives in this and similar licensing deals.
# Bottom line d-Matrix chose NVLink Fusion and MGX rack compatibility to simplify scaling and customer integration. The technical trade-offs emphasize memory bandwidth and capacity per rack to target large-model inference, while the licensing decision ties the company into Nvidia's data-center stack and economics.