Loading...
Loading...
Every large AI training run today depends on thousands of GPUs trading gradients back and forth, fast enough that the network never becomes the bottleneck. The technology making that possible — RDMA (Remote Direct Memory Access) — didn't start out built for AI at all. It started as a fix for a much older problem: the CPU getting in its own way.
RDMA's core trick is deceptively simple — it lets one server read and write directly into another server's memory, skipping the CPU and operating system kernel entirely. That single design choice cuts communication latency and CPU load dramatically, and it's why RDMA has become the foundation for high-performance computing (HPC), AI infrastructure, and financial high-frequency trading alike. Over the past 25 years, RDMA has moved from a dedicated, niche technology to a mainstream networking standard — and, more recently, to a technology with real domestic breakthroughs. Here's how it happened, in five stages.
In the 1990s, as high-performance computing scaled up, the standard TCP/IP protocol stack became the limiting factor: every packet meant another round of data copying and CPU cycles burned on network processing instead of computation. IBM, HP, and others began exploring ways to move data between machines without routing it through the CPU — the idea that would become RDMA.
In 1999, the IBTA (InfiniBand Trade Association) was founded, and in 2000 the InfiniBand specification was released with RDMA as its defining feature. That was the formal starting point for RDMA as a technology — though at this stage it lived entirely inside the InfiniBand ecosystem, reserved for high-end HPC.
Through the 2000s, InfiniBand kept improving — pushing latency below 1 microsecond and throughput steadily upward — which made it the default choice for top-tier HPC clusters. But dedicated InfiniBand hardware was expensive and locked customers into a closed ecosystem.
That cost pressure pushed the industry to ask a different question: could RDMA run on ordinary Ethernet? The RDMA Alliance formed in 2003, and in 2007 iWARP was standardized, bringing RDMA to standard Ethernet and enabling it to work across wide-area networks. At the same time, the OFA (OpenFabrics Alliance) built out the OFED software stack, making RDMA far more accessible to developers who weren't InfiniBand specialists.
This stage is where RDMA and Ethernet truly converged. RoCEv1 launched in 2010, running at Ethernet's Layer 2 and leaning on PFC (Priority Flow Control) for lossless transmission — effective, but unable to cross subnet boundaries. In 2014, RoCEv2 solved that: moving to Layer 3, adding cross-subnet routing, and layering in better congestion control. RoCEv2 struck the balance the market had been waiting for — near-InfiniBand performance at Ethernet-level cost — and it became the default RDMA protocol for data centers.
By this point, three distinct RDMA paths had emerged: InfiniBand, RoCE, and iWARP — each with a different cost/performance trade-off.
Cloud computing and AI training changed what RDMA was for. It moved out of specialized HPC clusters and into general-purpose data centers, becoming the interconnect underneath GPU cluster gradient synchronization, compute-storage separation in AI data centers, financial high-frequency trading, and distributed databases.
On the technical side, InfiniBand bandwidth climbed to 400G and then 800G, while RoCEv2 kept getting refined for larger-scale deployments. On the ecosystem side, Chinese vendors including Huawei and ZTE broke into a market that had long been dominated by a handful of overseas players, improving software compatibility along the way.
For most of RDMA's history, the core technology was controlled by a small number of overseas vendors. That's changed fast since 2025, as domestic vendors accelerated independent RDMA R&D. In 2026, the Sugon scaleX ten-thousand-card supercluster went live running scaleFabric — a natively developed RDMA technology built on fully self-developed IP and chips, delivering meaningful performance and cost advantages. Huawei, ZTE, and others have made parallel progress on RoCEv2 and domestic InfiniBand-equivalent architectures. Together, this is what a maturing, independently-controlled RDMA ecosystem looks like.
Zoom out across 25 years and the pattern is clear: RDMA has moved from closed and specialized to open and general-purpose, from overseas-controlled to domestically developed, and from single-use to multi-scenario. Looking ahead, expect continued performance gains, expansion into edge computing, a more mature domestic ecosystem, and RDMA playing an increasingly central role in the broader computing-power infrastructure buildout.
RDMA's value proposition — low latency, high throughput, lossless transmission, and full RoCEv2 compatibility — only shows up in production if the switches underneath it are built for it. AurCore's lineup covers the full RDMA network stack, from the 400G core backbone down to 25G/10G access, so intelligent computing centers, cloud data centers, and high-frequency trading environments can deploy RDMA end to end without stitching together mismatched hardware.
Built for the core backbone of high-end intelligent computing centers and supercomputer clusters, the AES8001 is purpose-fit for 400G-class RDMA traffic. It runs the full RoCEv2 protocol stack with hardware-based PFC and ECN, eliminating the packet loss and congestion jitter that undermine RDMA performance at scale — exactly what large AI model training and supercomputer-level parallel workloads demand. 32 high-density 400G QSFP-DD ports handle RDMA interconnection across large GPU clusters, while 2×10GE SFP+ ports keep management and low-speed service traffic separate. It's designed to slot into next-generation domestic RDMA architectures as the core switch for ten-thousand-card superclusters and large computing centers.
2. AES6201 — 100G RDMA Aggregation Switch (32×100G QSFP28)
The AES6201 sits at the data center aggregation layer and doubles as the core layer for small and mid-sized intelligent computing centers — the benchmark spec for mainstream 100G RDMA networking. Full wire-speed 100G ports carry RDMA traffic for distributed databases, compute-storage separation, and mid-sized AI training clusters, with full compatibility across both iWARP and RoCE. Its optimized packet-forwarding design cuts CPU scheduling overhead, staying true to RDMA's core promise of bypassing the kernel for efficient transmission. High-density 100G ports support horizontal scaling across multi-node server clusters, making the AES6201 a strong fit for mid-to-high-end HPC and financial high-frequency trading networks where cost and performance both matter.
Built for the access layer of data centers and edge computing RDMA deployments, the AES6102/AES6104 follows the classic 25G-access-to-100G-uplink architecture. 48×25G SFP28 ports connect batches of RDMA-enabled servers with the low latency dense deployments demand, while 8×100G uplinks provide non-blocking connectivity to 100G/400G core devices for end-to-end lossless RDMA. Built-in visualized buffer optimization and intelligent congestion control make it a precise fit for edge computing and mid-sized distributed storage clusters — solving the latency and packet-loss problems that hold traditional networks back, and supporting RDMA's push into general-purpose data centers.
For teams that need flexible, cost-effective RDMA access, the AES6006 targets general data centers and campus computing environments. 48×10G Base-T electrical ports connect directly to standard servers and endpoints, enabling RDMA deployment without optical modules — a real cut in wiring and hardware costs. 6×100G QSFP28 uplinks provide enough headroom for concurrent RDMA traffic from multiple endpoints. With full RDMA protocol compatibility and rapid lossless-network deployment, the AES6006 is built for cost-sensitive, large-scale RDMA rollouts — the switch that takes RDMA from high-end use cases into everyday commercial deployment.