RoCEv2 · DLB · Ultra Ethernet

InfiniBand vs Ethernet for AI Fabrics

For most AI fabrics, Ethernet. For latency-critical HPC, InfiniBand. Modern Ethernet with RoCEv2 now runs production fabrics at 400G and 800G on open, multi-vendor hardware, and OcNOS-DC delivers it today. This guide covers where each one still fits.

Two-thirdsEthernet share, AI scale-out
up to 800Gproduction RoCEv2 fabrics
1 NOSOcNOS-DC on open hardware
600+operator networks on OcNOS
Two fabrics, two operating models

The same GPUs, two very different networks

On the left, a single-vendor InfiniBand fabric: one silicon vendor, one switch line, one NIC ecosystem, and a subnet manager separate from the rest of the data center. On the right, a multi-vendor open Ethernet fabric: RoCEv2 or UEC NICs from any vendor, Broadcom switch silicon, OcNOS-DC as the NOS, and the same protocols you already run.

InfiniBand single-vendor fabric vs multi-vendor Ethernet fabric Two side-by-side fabric topologies. Left: a single-vendor InfiniBand fabric with two IB switches and three GPUs. Right: a multi-vendor Ethernet fabric with two OcNOS-DC spines and three GPUs with RoCEv2 or UEC NICs. Bottom labels contrast a single-vendor stack with an open multi-vendor stack. INFINIBAND · SINGLE-VENDOR ETHERNET · OPEN MULTI-VENDOR IB Switch-1Quantum-class IB Switch-2Quantum-class GPUIB NIC GPUIB NIC GPUIB NIC SINGLE NIC + SWITCH VENDOR SUBNET MANAGER · LIMITED MULTI-TENANCY Spine-1OcNOS-DC · TH5 Spine-2OcNOS-DC · TH5 GPURoCEv2/UEC GPURoCEv2/UEC GPURoCEv2/UEC MULTI-VENDOR · OPEN HARDWARE RoCEv2 · DLB · GLB · UEC · EVPN-VXLAN · gNMI SINGLE-VENDOR STACK vs OPEN MULTI-VENDOR STACK
How they compare

InfiniBand and Ethernet, axis by axis

InfiniBand was purpose-built for low-latency, lossless RDMA, and for two decades that gave it a real edge for tightly coupled HPC. Modern Ethernet, built on the DCB stack, RoCEv2, and increasingly DLB and UEC, has spent the last several years closing that gap. Which gaps still matter depends on the workload.

Axis InfiniBandsingle-vendor EthernetRoCEv2 / UEC, open
Latency floorVery low end-to-end NIC-to-NIC; typical hundreds of nanoseconds switch hop.Higher floor than IB by hundreds of nanoseconds, but well below the threshold that affects most distributed-training collectives at scale.
Loss toleranceLossless by architecture (credit-based flow control).Lossless via PFC + ECN + DCQCN. Production-grade today; UEC further reduces dependence on PFC pause.
Multi-path / load balancingAdaptive routing built into the spec.Static ECMP, plus DLB for adaptive single-hop, GLB (OcNOS 7.1) for end-to-end, UEC packet-spray for next-gen.
Vendor ecosystemEffectively single-vendor for both NIC and switch silicon.Multi-vendor at every layer: ASIC, switch, NIC, NOS, optics. UEC is explicitly designed for vendor-neutral interop.
Operational modelSubnet manager (UFM-class). Different from rest of DC. Separate skills, separate tooling.Same BGP, EVPN, gNMI you already run. Same automation tools (Ansible, NETCONF, OpenConfig) as the rest of DC.
Multi-tenancyLimited; partitioning exists but is not a first-class concept.First-class via EVPN-VXLAN. GPU-as-a-Service, multi-team clusters, shared infra all natural.
Long-haul DCINot designed for it; needs IB-over-WAN gateways.Native via 400G ZR/ZR+ coherent pluggables and EVPN inter-DC.
Storage convergenceStorage runs alongside compute; needs IB-attached storage.NVMe-oF, NFS, S3 all over the same Ethernet fabric.
Cost / port (typical 400G+)Premium; single-vendor pricing.Open-hardware spine + OcNOS-DC NOS materially undercuts vendor-locked alternatives.
Roadmap velocityDriven by one vendor's release cadence.UEC consortium (AMD, Arista, Broadcom, Cisco, HPE, Intel, Meta, Microsoft, Oracle) drives openly published spec evolution.
The decision

Where each one wins

The right answer is workload-specific. Pick InfiniBand where an absolute latency floor is the requirement; pick Ethernet where the operating model, cost, or reach across sites carries the decision.

Pick InfiniBand when

The latency floor is contractual

HPC simulation where the absolute latency floor matters more than total cost of ownership, and tight, captive single-tenant clusters where lock-in is acceptable.

Pick Ethernet when

The operating model matters

Multi-tenant GPU-as-a-Service and clusters that share infrastructure with the rest of the data center, where one operating model, one tooling stack, and a multi-vendor supply chain win.

Pick Ethernet when

Cost per GPU-flop is the gate

Open-hardware spines with OcNOS-DC remove the single-vendor network premium. On a multi-thousand-GPU cluster, that is a material share of the hardware budget.

Pick Ethernet when

The fabric spans data centers

If a training run will ever cross two halls or two regions, coherent DCI, EVPN inter-DC, and standard multi-vendor optics make it a one-day problem rather than a quarter-long line-system project.

Where the gap closed

What modern Ethernet added

Three developments moved production AI fabrics onto Ethernet: genuinely lossless behaviour, adaptive routing that keeps flows off congested uplinks, and a spray-friendly next-generation transport.

Lossless behaviour

RoCEv2 with PFC, DCQCN, and the PFC deadlock watchdog delivers the lossless RDMA transport AI collectives need, production-grade today on standard Ethernet.

Adaptive routing

Static ECMP collisions on AI workloads are real, but DLB rebins flowlets on local congestion in sub-millisecond windows, and GLB in OcNOS 7.1 extends that to end-to-end path scoring.

Spray-friendly transport

Ultra Ethernet (UEC 1.0, June 2025) brings packet spray, multi-path RDMA, and out-of-order delivery to standard Ethernet. Build on RoCEv2 now and keep a clean path to UEC as NICs ship.

The TCO conversation

The network is a small line item, so spend the comparison on what it frees up

Over five years the switching layer is a single-digit percentage of cluster TCO, rising once NICs, optics, and cabling are counted. The useful question is not the line item, but what the saved capital funds.

  • For equivalent capacity, single-vendor InfiniBand carries a price premium over open-hardware Ethernet.
  • The saved capital funds more GPUs, a larger storage tier, or a second site for resilience.
  • Where the fabric is multi-tenant or shared with the rest of the data center, one network model is worth more than the line-item difference.
The IP Infusion view

Both have a place, and most AI fabrics belong on Ethernet

Ethernet does not win every workload, but the operational and economic case is decisive for most, and it strengthens as the technical gap closes.

Both have a place

Ethernet does not win every workload. Tight HPC clusters with absolute-floor latency requirements still favor InfiniBand.

Most AI fabrics belong on Ethernet

Production AI training and inference at hyperscale is moving to Ethernet as the operational and economic case grows and the technical gap keeps closing.

OcNOS-DC is the open path

RoCEv2 today, DLB today, GLB next, UEC as NICs ship. One NOS, one feature roadmap, on validated open hardware from Edgecore, UfiSpace, and others.

FAQ

InfiniBand vs Ethernet, answered

Is Ethernet fast enough for AI training versus InfiniBand?
For most distributed-training collectives at scale, yes. Ethernet has a higher latency floor than InfiniBand by hundreds of nanoseconds, but that sits below the threshold that affects most workloads once RoCEv2 with PFC, DCQCN, and DLB is configured correctly.
When should I still pick InfiniBand?
Choose InfiniBand for tight HPC simulation where the absolute latency floor matters more than total cost of ownership, and for captive single-tenant clusters where single-vendor lock-in is acceptable.
When does Ethernet win for an AI fabric?
Ethernet wins for multi-tenant GPU-as-a-Service, clusters that share one operational model and tooling with the rest of the data center, cost-sensitive multi-thousand-GPU builds, and fabrics that extend across data centers via coherent DCI and EVPN.
How does OcNOS close the gap with InfiniBand?
OcNOS-DC delivers RoCEv2 lossless transport, DLB adaptive routing, and the PFC deadlock watchdog today, GLB in 7.1 for end-to-end path scoring, and UEC support as NICs ship, all on validated open hardware from vendors such as Edgecore and UfiSpace.
Are hyperscalers replacing InfiniBand with Ethernet for AI?
Largely yes, though InfiniBand is not going away. According to Dell'Oro Group, Ethernet passed InfiniBand in AI back-end (scale-out) switching and reached about two-thirds of the market by early 2026, up from InfiniBand's roughly 80 percent share in late 2023. Dell'Oro also notes a strong InfiniBand rebound on the NVIDIA Blackwell Ultra 800G ramp, so it remains the specialist choice for tightly coupled HPC. Meta has described training its largest models over a RoCE Ethernet fabric on one of two 24,576-GPU clusters. The Ultra Ethernet Consortium, whose steering members include AMD, Arista, Broadcom, Cisco, Eviden, HPE, Intel, Meta, Microsoft and Oracle, standardizes this direction. The reasons are operational and economic: one network model shared with the rest of the data center, a multi-vendor supply chain, and lower cost per port at 400G and 800G.
What is the best alternative to InfiniBand for AI cluster networking?
Ethernet with RoCEv2 is the mainstream alternative. It carries RDMA losslessly using PFC, ECN and DCQCN, adds adaptive routing through DLB, and gains packet spray and multi-path RDMA through Ultra Ethernet as UEC NICs ship. On open hardware running OcNOS-DC, it delivers this on a multi-vendor stack rather than a single-vendor fabric.
What is the difference between Ultra Ethernet and InfiniBand?
Ultra Ethernet (UEC) brings the transport techniques that defined InfiniBand to standard Ethernet: packet spray across all paths, multi-path RDMA, out-of-order delivery, and selective retransmission. The UEC 1.0 specification, published in June 2025, defines this transport. The difference is the ecosystem: InfiniBand is effectively single-vendor for NIC and switch silicon, while Ultra Ethernet is an open, multi-vendor specification any vendor can build to.
How much does an Ethernet AI fabric cost compared to InfiniBand?
The switching layer is a modest share of AI cluster total cost of ownership, usually a single-digit percentage over five years, though the figure rises once NICs, optics, and cabling are counted. For equivalent capacity, single-vendor InfiniBand generally carries a price premium over open-hardware Ethernet running OcNOS-DC. The more useful question is what the saved capital funds, whether that is more GPUs, a larger storage tier, or a second site for resilience.

Sizing your next cluster? Get a workload-specific review

Tell us the workload and the GPU scale, and an IP Infusion engineer will run the maths with you, or start with a first-pass leaf-spine layout in the AI Fabric Design Suite.