Theregister iconTheregisterSep 10, 2026 ~7 min source read

d-Matrix adds NVLink Fusion and MGX rack support to scale Raptor inference systems

Startup d-Matrix will integrate Nvidia’s NVLink Fusion interconnect and MGX rack reference designs into upcoming Raptor accelerators, enabling racks with up to 144 accelerators on a single all-to-all NVLink fabric and compatibility with Nvidia networking and rack components.

d-Matrix drinks the Nvidia Kool-Aid with NVLink Fusion and MGX rack designs

Share this story

Send the public story page.

Useful takeaways from this story.

d-Matrix will support NVLink Fusion and MGX rack designs so customers can deploy its Raptor accelerators in Nvidia-compatible racks and NVSwitch fabrics.

Planned NVL144 racks will host up to 144 Raptor-based XPUs, delivering around 2.3 TB of memory capacity and roughly 7.2 PB/s aggregate memory bandwidth, aimed at large inference workloads.

Raptor cards use 3D-stacked DRAM to prioritize memory bandwidth (up to ~100 TB/s on full Raptor), trading off raw capacity for SRAM-like bandwidth to reduce chips needed per model.

# What happened AI-infrastructure startup d-Matrix will license Nvidia's NVLink Fusion interconnect and adopt MGX rack-scale reference designs for its next-generation Raptor accelerators. The move lets d-Matrix build systems that plug into the same rack and NVSwitch fabrics used by Nvidia, and to deliver racks with up to 144 accelerators tied together by a single all-to-all NVLink fabric.

# Why it matters Scaling an accelerator architecture across large clusters requires a high-speed, low-latency fabric. By using NVLink Fusion and MGX rack designs, d-Matrix avoids designing its own rack interconnects and leverages an established ecosystem of NVSwitch appliances, Vera CPUs, BlueField and ConnectX NICs, and SpectrumX Ethernet. That reduces integration work for customers who already deploy Nvidia-compatible racks.

# Technical snapshot Raptor accelerators use bonded compute logic on top of 3D-stacked DRAM to push memory bandwidth close to SRAM-like levels. At Hot Chips the company showed a full Raptor card with 32 GB of 3D-stacked DRAM capable of about 100 TB/s of memory bandwidth. d-Matrix's NVL144 rack concept uses smaller XPUs—roughly half the size of the Hot Chips Raptor card—with about 16 GB of 3D-DRAM and ~50 TB/s of per-XPU bandwidth.

With 144 XPUs per rack, the design targets approximately 2.3 TB of total memory capacity and a peak aggregate memory bandwidth near 7.2 petabytes per second. That memory profile is pitched at inference for very large models: d-Matrix estimates the NVL144 configuration can serve models exceeding four trillion parameters at 4-bit precision.

# How this fits inference workflows

# Business and ecosystem effects

# What to watch next

  • d-Matrix timeline for shipping NVL144 systems and availability of the full Raptor cards.
  • Real-world throughput and latency metrics for heterogeneous GPU+Raptor inference pipelines.
  • Any commercial or investment terms disclosed that clarify Nvidia's incentives in this and similar licensing deals.

# Bottom line d-Matrix chose NVLink Fusion and MGX rack compatibility to simplify scaling and customer integration. The technical trade-offs emphasize memory bandwidth and capacity per rack to target large-model inference, while the licensing decision ties the company into Nvidia's data-center stack and economics.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app