RoCEv2 · DCQCN · DLB · UEC 1.0

An open 800G AI fabric on standard Ethernet

You can build a production GPU-cluster fabric on standard Ethernet and keep your NIC, GPU, and switch choices open. IP Infusion delivers that fabric complete: validated 800G open switches from Edgecore or UfiSpace running OcNOS-DC, supported under one contract. It carries lossless RoCEv2 RDMA with PFC, ECN, DCQCN, and sub-millisecond DLB, and aligns with the Ultra Ethernet 1.0 fabric profile.

600+Production OcNOS networks
60+Countries in service
26 yrRouting stack in production
Multi-
thousand
GPU reference designs
The stakes

What moves job completion time

What matters at scale is how fast jobs finish and how busy the GPUs stay, not switch throughput. Every time the cluster pauses to synchronize, idle GPUs waste capacity, so the fabric has to drop nothing and react to congestion the moment it starts.

OcNOS-DC exposes every setting, so your team tunes the fabric against your real GPU collective traffic (xCCL: NCCL, RCCL, oneCCL) instead of accepting a vendor's fixed profile. Each pattern below is one way a cluster stalls, and how OcNOS-DC keeps it moving.

Training sync (AllReduce)
Every GPU talks to every other GPU at once
Standard load balancing pins these large flows to one uplink, so some links jam while others sit idle and the sync waits on the slowest one.
DLB shifts flows to less-busy paths in under a millisecond.
GLB is on the OcNOS roadmap to balance across the whole fabric, not just the local hop.
Result: no traffic hot spots; the sync step runs at close to full speed.
Traffic bursts (incast)
Many senders hit one port in microseconds
A dropped packet restarts the whole collective, and pausing too hard freezes the link, so the fabric has to head congestion off early.
DCQCN slows senders early, before anything overflows.
PFC watchdog clears a stuck port on its own.
Result: jobs ride out bursts, and a stalled port recovers without a manual reset.
Scale-out (multi-rail)
One flow needs every parallel path at once
Today one flow rides a single path, leaving the other rails unused.
Ultra Ethernet (UEC 1.0) spreads a single flow across every path at once.
→ The switch you buy today carries over when UEC NICs arrive, with no swap.
Result: the slowest transfers speed up as UEC NICs roll out.
~55% → 90%+

Reference benchmark. DLB lifts fabric utilization from about 55% with standard load balancing to 90%+ on the same switches, with no extra uplinks. (Industry-published Broadcom figure for Tomahawk 4 and 5.)

DLB deep-dive →
Reference topology

800G spine-leaf, lossless end to end

This is the leaf-spine (Clos) design your team already knows, built to run GPU traffic without loss: a routed eBGP underlay spreads traffic evenly across every path, lossless priority queues protect the RoCEv2 flows, and a separate management network brings switches up and streams telemetry on its own. Leaf-attached NVMe-oF and NFS storage for checkpoints and datasets sits adjacent to the GPU racks; see the DC Fabric for storage-fabric detail. Hover any node for platform, port count, and chip.

AI fabric topology: 800G spine-leaf with parallel leaves, full-mesh eBGP, an isolated management bus, and leaf-attached storage
AI Fabric: 800G spine-leaf, full-mesh eBGP, an isolated management bus, and leaf-attached storage.
The hardware

Validated open switches

OcNOS-DC runs on Broadcom Tomahawk and Trident switches from Edgecore and UfiSpace, so every platform is validated, orderable, and second-sourceable.

Edgecore AIS800-64D 800G spine switch
Spine · Tomahawk 5

Edgecore AIS800-64D

64 x 800G · 51.2 Tbps

QSFP-DD800 · breakout to 2x400G / 4x200G / 8x100G

Datasheet →
UfiSpace S9321-64E 800G spine switch
Spine · Tomahawk 5

UfiSpace S9321-64E

64 x 800G · 51.2 Tbps

QSFP-DD800 · breakout to 2x400G / 4x200G / 8x100G

Datasheet →
Edgecore AS9736-64D 400G leaf switch
Leaf · Tomahawk 4

Edgecore AS9736-64D

64 x 400G · 25.6 Tbps

QSFP56-DD 400G · breakout to 4x100G / 4x25G · DAC / AOC

Datasheet →

See all 40+ validated platforms in the HCL →

One vendor behind the software, the hardware, and support

SINGLE-VENDOR SUPPORT

One contract, one team to call

A single IP Infusion contract covers the OcNOS-DC software and the validated switch hardware together, so there is one TAC and one SLA instead of a separate NOS and hardware vendor. A 24/7 support portal is available on the Premium and Enterprise tiers.

See support details
IP MAESTRO MANAGEMENT

One console for the fabric

IP Maestro is an element management system for OcNOS: topology, faults, configuration, and software images across every switch over NETCONF, sitting beneath your existing OSS. Manages large multi-site fleets from one console.

Explore IP Maestro
AUTOMATED DAY 2

Streaming telemetry and zero-touch

gNMI streaming telemetry feeds Prometheus and Grafana, DCBX pushes the correct lossless settings to each server automatically, and zero-touch provisioning brings switches up configured at boot.

Explore automation
Open alternative to Spectrum-X

The open, multi-vendor alternative to NVIDIA Spectrum-X

OcNOS-DC runs a lossless RoCEv2 AI fabric on standard Ethernet switches from more than one open-hardware vendor, under one support contract. It matches the fabric mechanisms an AI training job depends on, while keeping the NIC, GPU, and hardware choices open instead of locked to a single vendor.

Open AI fabric on OcNOS-DC versus a closed single-vendor AI stack. Last verified: Jul 2026.
What decides the deal Open AI fabric (OcNOS-DC on Edgecore / UfiSpace) Closed single-vendor stack (NVIDIA Spectrum-X)
Hardware sourcing Open merchant silicon from more than one vendor (Edgecore, UfiSpace) Single-vendor switch, NIC, and GPU
NIC and GPU choice Open; swap NICs and GPUs independently Tied to the NVIDIA NIC and GPU roadmap
Lossless RoCEv2 (PFC, ECN, DCQCN) details →
Adaptive load balancing sub-millisecond DLB flowlet rebinding details → (proprietary)
PFC deadlock protection deadlock watchdog with auto-drain details →
Ultra Ethernet (UEC 1.0) tracks the UEC 1.0 fabric profile details → Aligned; also a UEC member
Streaming telemetry and automation gNMI, OpenConfig, ZTP, DCBX, Ansible (vendor pipeline)
Single support contract (switch and software) one TAC, one SLA (single vendor)
GPU-silicon back-end integration Standard Ethernet; no GPU-silicon co-design Advantage NVIDIA: NIC, switch, and GPU co-designed
Fabric-controller layer Managed through gNMI telemetry and your OSS or IP Maestro Advantage NVIDIA: integrated fabric-controller layer
UEC certification Aligned with the UEC 1.0 fabric profile (not a certification claim) NVIDIA is a UEC steering member

NVIDIA, Spectrum-X, Quantum, and ConnectX are trademarks of NVIDIA Corporation. IP Infusion is not affiliated with and does not endorse NVIDIA; the comparison reflects OcNOS-DC capabilities verifiable in the feature matrix.

The 2026 AI fabric landscape

Where OcNOS-DC sits

By 2026 almost every option clears the same technical bar: lossless RoCEv2, congestion control, adaptive routing, Ultra Ethernet alignment. So the decision comes down to the shape of the deal: an open operating system or a locked stack, open or locked hardware, standard Ethernet or closed InfiniBand. Here is where each one leaves you.

Solution shape Examples Trade-off
Open NOS, AI-hardened, UEC-aligned OcNOS-DC on Edgecore / UfiSpace Same Broadcom silicon, the same technical floor. DCQCN tuned for RDMA collective traffic, sub-ms DLB, GLB on the OcNOS roadmap, PFC deadlock watchdog, UEC 1.0 fabric profile. Single-vendor support. No NIC, GPU, or hardware lock-in.
Closed vertical AI stack NVIDIA Spectrum-X + Quantum + ConnectX Excellent integrated performance. NIC, switch, and fabric software locked to one vendor, and to one GPU roadmap.
Locked merchant-silicon NOS Arista EOS · Cisco NX-OS · Juniper Junos Same Broadcom silicon underneath. Per-port licensing premium. Telemetry and tuning constrained to the vendor's own pipeline.
Cell-based proprietary chassis fabric DriveNets Network Cloud Different architecture: scheduled cell fabric, not Ethernet NOS. Strong at hyperscale; not portable to standard switches.
Closed-loop InfiniBand NVIDIA Quantum InfiniBand Lowest latency for tight collectives today. Separate cabling, separate operations, single-vendor ecosystem. UEC closes the gap on Ethernet.
Open NOS, no AI hardening Community SONiC Open hardware, free software, no SLA. RDMA-tuned DCQCN defaults, deadlock watchdog, and tuning maturity are left entirely to the operator.

Each option targets a different priority. OcNOS-DC leads with open hardware, a complete solution, and no lock-in.

AI Fabric size guide

Size your GPU fabric in minutes

Enter your GPU count and NIC speed. The AI Fabric Design Suite maps it to a leaf-spine topology, switch and port counts, and a bill of materials you can take straight to a quote. No spreadsheet required.

Common questions

Frequently asked questions

Is OcNOS-DC actually "AI-native", or just RoCEv2 with extras?
No merchant-silicon Ethernet NOS is literally AI-native: none reason about xCCL (NCCL / RCCL / oneCCL) collectives or schedule jobs at the switch; that lives in the NIC and scheduler. OcNOS-DC implements every fabric mechanism a 2026 AI workload needs, lossless RoCEv2, DCQCN buffer profiles tuned for RDMA collective traffic patterns, sub-millisecond DLB flowlet rebinding, PFC deadlock watchdog, UEC 1.0 alignment, and stays out of the layers above. "AI-aware fabric" usually just means one vendor sells NIC + switch + scheduler as one locked SKU.
Where does OcNOS-DC stop, and where do the NIC and cluster scheduler take over?
OcNOS-DC owns layer 1: lossless RDMA transport, congestion control, adaptive routing, deadlock recovery, telemetry. The NIC owns layer 2 (xCCL, RDMA verbs, packet spray, GPU-direct memory); the scheduler owns layer 3 (job placement, gradient-sync windows, tenant isolation). OcNOS-DC streams gNMI telemetry into layer 3 but never tries to be the scheduler. That separation keeps your NIC, GPU, and orchestration swappable.
How does OcNOS AI Fabric compare to NVIDIA Spectrum-X, SONiC, Arista, Cisco, or DriveNets?
Spectrum-X is a closed NVIDIA NIC + switch + software stack: excellent performance, single-vendor lock-in. Arista, Cisco, and Juniper run similar RoCEv2 features on locked hardware with proprietary licensing. Community SONiC is open but ships no AI-hardened defaults, watchdog, or SLA. DriveNets DDC is a proprietary cell fabric, not an Ethernet NOS. OcNOS-DC: open NOS on the same Broadcom silicon, UEC-aligned, DCQCN tuned for RDMA collective traffic, 24/7 SLA, same technical floor, no lock-in.
What does Ultra Ethernet (UEC) 1.0 mean for OcNOS AI Fabric?
The switch you buy today should carry forward when Ultra Ethernet NICs arrive, and on OcNOS-DC it does. Your fabric runs fully supported RoCEv2 plus DCQCN plus DLB in production now, and OcNOS-DC tracks the UEC 1.0 fabric profile, which parallelizes each flow across every path instead of pinning it to one ECMP hash, so you move to UEC NICs with no NOS or hardware swap. OcNOS-DC is aligned with the UEC 1.0 fabric profile; alignment is not a certification claim. See the Ultra Ethernet deep-dive.
What is RoCEv2 and why does it require a lossless Ethernet fabric?
Your training collectives, AllReduce and AllGather, move data GPU-to-GPU over RoCEv2 with no CPU in the path, and RDMA never retransmits, so a single dropped packet restarts the operation across every GPU in the job. That is why production RoCEv2 needs a genuinely lossless fabric (PFC plus ECN), and OcNOS-DC ships the RoCEv2 buffer profiles and DCQCN defaults tuned for RDMA collective traffic so your team is not building that lossless behavior from scratch.
How does OcNOS-DC keep the fabric lossless, and what protects against PFC deadlock?
Three mechanisms: PFC pauses per-priority traffic before buffers overflow, ECN marks packets early to slow senders, and ETS keeps RDMA flows ahead of lower-priority traffic. On top, a per-port, per-priority deadlock watchdog detects paused-queue cycles and auto-drains the queue before jobs hang: the failure mode that used to force mid-job switch power-cycles. PFC over L3 is supported across routed boundaries.
What is DLB, and what is GLB on the OcNOS roadmap?
During AllReduce your biggest flows collide when standard ECMP pins each one to a single uplink for its lifetime, and the sync waits on the jammed link. DLB reads live ASIC queue-depth telemetry and rebinds flowlets to less-loaded paths in under a millisecond, so it recovers that lost throughput at the local hop today. GLB is on the OcNOS 7.x roadmap to extend the same idea fabric-wide: spines publish path-quality telemetry back to the ingress leaves so routing scores the full multi-hop path across large multi-thousand-GPU clusters.
What scale does OcNOS AI Fabric support, and what are the validated reference designs?
OcNOS-DC supports 400G and 800G leaf-spine fabrics. Tomahawk 5 spines (Edgecore AIS800-64D, UfiSpace S9321-64E) deliver 51.2 Tbps / 64 × 800G; Tomahawk 4 leaves run 400G / 25.6 Tbps with on-chip buffering; Trident 4 covers smaller 100G/400G fabrics. Reference designs cover rail-only, rail-optimized, and 3-stage Clos topologies for large multi-thousand-GPU clusters: see the AI Fabric Topologies deep-dive.
Does OcNOS-DC support automation and telemetry for AI fabric operations?
Yes. DCBX automates server-to-switch RoCEv2 config, ZTP (IPv4/IPv6) handles zero-touch onboarding, and gNMI streams on-change telemetry over OpenConfig YANG. PFC pauses, ECN marking, DCQCN thresholds, and buffer depths are gNMI sensor paths consumable by Prometheus, InfluxDB, Telegraf, Grafana, or any OpenTelemetry pipeline. Ansible playbooks cover Day-0 through Day-2, with a Terraform provider on the roadmap. IP Maestro, an element management system for OcNOS, manages topology, faults, configuration, and software images over NETCONF, and manages large multi-site fleets from one console.
What switch hardware and 800G optics does the OcNOS AI fabric support?
The fabric runs on validated Broadcom Tomahawk 4 and 5 switches from Edgecore and UfiSpace. The 800G spines (Edgecore AIS800-64D and UfiSpace S9321-64E, Tomahawk 5, 64 x 800G QSFP-DD800, 51.2 Tbps) break out to 400G, 200G, and 100G; the 400G leaf (Edgecore AS9736-64D, Tomahawk 4, 64 x 400G QSFP56-DD, 25.6 Tbps with on-chip buffering) breaks out to 100G and 25G. Optics have proven interoperability from multiple vendors, so contact us for the transceiver list for your build. The switch platforms are listed in the HCL.
Who provides support for the AI fabric, and is the hardware covered too?
One IP Infusion contract covers the OcNOS-DC software and the validated switch hardware together, so there is one TAC and one SLA. Support is tiered: Standard includes email TAC, business-hours phone, and hardware RMA coordination with an 8-hour P1 response; Premium adds a 24/7 customer support portal and a 2-hour P1 response; Enterprise adds a 30-minute P1 response and a Customer Success Manager. The 24/7 portal is available on Premium and Enterprise.