AI fabric design tool

Design a non-blocking RoCEv2 GPU fabric end to end, vendor-neutral. Size the leaf and spine switches, plan the transceiver and cable counts, compare InfiniBand against Ethernet for your cluster, and get a lossless congestion profile. This designs the network only. It is not a cost estimate or a bill of materials.

Step 1

Size the GPU fabric

Pick a starting size, or enter your own GPU count, port speed, and switch silicon. The tool sizes a folded two-tier leaf-spine pod and flags when a cluster crosses into three-tier scale.

SPINE Spine 1Spine 2Spine 3 LEAF Leaf 1Leaf 2Leaf 3Leaf 4 GPU SERVERS
leaf switches
spine switches
total switches

Counts assume a folded leaf-spine (Clos) fabric with one fabric NIC per GPU, where each leaf splits its ports between GPUs and spine uplinks. Rail-optimized wiring (multiple NICs per GPU) changes traffic locality, not these switch counts. See the AI fabric topology reference designs for the rail-optimized and three-tier layouts.

Step 2

Component summary

Fabric link, transceiver, and cable quantities for the sized design. These are counts for planning only, not a bill of materials, quote, or pricing.

ComponentBasisQuantity
GPU-facing linksGPUs × 1 NIC
Leaf to spine fabric linksleaves × uplinks
Fabric transceiverslinks × 2 ends
Fabric cables1 per link

Media by reach, chosen per your cabling: up to 3 m DAC or AEC, up to 50 to 100 m AOC or SR, up to 500 m DR (the workhorse), up to 2 km FR. Reach classes only. No pricing, SKUs, or vendor part numbers. On a three-tier design, add the spine to super-spine links on top of the counts above.

Decision

InfiniBand vs Ethernet

A vendor-neutral scorecard across six factors. It shows where each factor leans for your inputs and why. There is no single winner here by design. Use it as an input to a design review, not a final recommendation.

FactorInfiniBandEthernet / RoCEv2Ultra Ethernet (trajectory)
Peak performanceleans InfiniBandA mid-teens percent raw throughput edge at low tuning effort, plus SHARP in-network reduction that offloads collective operations from the GPUs.Matches InfiniBand with proper tuning. A 24,000-GPU RoCEv2 training cluster has run in production at performance parity. The gap is tuning effort, not the ceiling.Adds packet spray and NIC-side reordering to cut tail latency and reduce PFC dependence, narrowing the out-of-box gap.
Costleans EthernetCarries a meaningful premium at fabric scale, commonly cited in the 30 to 60 percent range versus comparable Ethernet.Merchant-silicon switches and a broad optics supply chain generally make it the lower-cost fabric per port.Same open merchant-silicon economics as Ethernet.
Opennessleans EthernetSingle-vendor stack end to end (switches, NICs, manager), so multi-vendor sourcing is limited.Open and multi-vendor: standards-based switches, NICs, and NOS from many suppliers, preserving supply-chain optionality.Industry consortium standard (UEC 1.0, 2025) reinforcing the open direction.
Operational simplicityneutralUFM gives a single, purpose-built fabric manager, which is turnkey for teams without deep RDMA staff.Rides the large Ethernet talent pool and standard tooling, but lossless RoCEv2 needs deliberate end-to-end tuning.Aims to reduce the tuning burden, moving Ethernet toward IB-like out-of-box behavior.
Multi-tenancy and convergenceleans EthernetTypically a dedicated back-end island, without native multi-tenant isolation or storage convergence.One converged Ethernet fabric can carry back-end, storage, and front-end traffic with standard tenant isolation.Extends the converged-Ethernet direction on one open fabric.
Trajectoryleans EthernetMature and stable, performance-leading at low tuning, but on a single-vendor roadmap.Ethernet crossed over to lead AI back-end deployments by mid-2025, with broad ecosystem momentum.UEC 1.0 (2025) is a multi-vendor spec explicitly targeting AI and HPC back-end fabrics.

Assumptions: performance is compared with tuning (untuned Ethernet trails, well-tuned Ethernet reaches parity); cost is a range, never a dollar figure; openness means multi-vendor sourcing, not a quality judgment; Ultra Ethernet reflects a 2025-onward direction, not shipping product. Collective tuning refers to xCCL libraries generically. This informs the decision; validate a fabric choice with an IP Infusion engineer.

Profile

RoCEv2 lossless profile

RoCEv2 is not a single switch feature. It is a profile composed from PFC, ECN, and ETS, with DCQCN doing the steady-state work and PFC as a backstop. This recommends the mechanisms and the order to apply them for your design. It does not emit threshold values and it does not touch any device.

MechanismRole in the lossless profileOcNOS availability
PFCNo-drop guarantee for the RoCE class. Last-resort backstop, not the primary control.All 4 AI platforms
ECNMarks congestion early so senders throttle before queues overflow. Tune its marking point below the PFC trigger.All 4 AI platforms
DCQCN (ECN + PFC)Primary steady-state control. ECN marking drives the DCQCN reaction at the NICs; PFC backs it up. Composed, not a standalone feature.All 4 AI platforms
ETSBandwidth guarantees and class isolation so the lossless class is not starved by best-effort traffic.All 4 AI platforms
PFC-deadlock watchdogDetects and breaks PFC pause deadlock or storm conditions, the safety net for the PFC backstop.All 4, OcNOS 7.0
DLB (Dynamic Load Balancing)Spreads elephant RoCE flows that collide under static ECMP. Lifts fabric utilization from roughly 55 percent toward 90 percent and above.All 4 AI platforms
Dynamic ECNAdapts ECN marking to live queue conditions, reducing manual re-tuning as load shifts.TH5 only, OcNOS 7.0
DLB reactive / randomAdvanced DLB placement modes for tighter flow spreading. On TH4, use standard DLB.TH5 only, OcNOS 7.0
GLB (Global Load Balancing)Fabric-wide load balancing beyond local DLB decisions.Roadmap, OcNOS 7.1 train
Ultra EthernetPacket spray with NIC-side reordering and reduced PFC dependence, the trajectory for lossless AI Ethernet.Roadmap, UEC profile

Advisory only. These recommendations are not pushed to any device and no threshold values are generated here. Availability is per the OcNOS-DC Feature Matrix and Hardware Compatibility List. Validate against your selected platforms and OcNOS release before deployment.

Estimate

Fabric power

A typical power range for the fabric switches in the sized design. Ranges only, optics-loaded. Network switches are a small part of total cluster power; the GPUs dominate.

kW typical low
kW loaded high

Cooling: rear-door or direct-liquid cooling is typically considered above roughly 30 to 40 kW per rack, driven by GPU-server density, not the network. This is guidance, not a computed thermal load. Switch power ranges are typical values for 51.2T (Tomahawk 5) and 25.6T (Tomahawk 4) class boxes with optics populated; confirm exact figures on each platform datasheet.

Matching OcNOS-DC platforms

The fabric runs a single OcNOS-DC image on open Broadcom Tomahawk hardware. These are the supported 800G and 400G platforms it designs onto.

This tool provides a structural network-design estimate for planning only. It is not a performance guarantee, a bill of materials, or a cost estimate. Actual designs depend on rail optimization, NIC count per GPU, cabling, and failure-domain choices. GPU counts are reference-design ceilings from Clos radix math, not measured figures. Broadcom and Tomahawk are trademarks of Broadcom Inc.; other names are the trademarks of their respective owners.

Resource

AI Fabric Reference Design

Get the reference design PDF: topology, switch counts, component summary, and a RoCEv2 lossless starting profile.

Download PDF
FAQ

Frequently asked questions

How do I size a leaf-spine fabric for a GPU cluster?
In a non-blocking two-tier leaf-spine fabric, each leaf switch dedicates half its ports to GPUs and half to spine uplinks. The number of leaf switches is the GPU count divided by the GPU-facing ports per leaf, and the number of spine switches is the fewest needed to carry every leaf uplink without oversubscription. A single-leaf cluster needs no spine tier. This tool computes those counts for Broadcom Tomahawk 4 and Tomahawk 5 at 400G or 800G.
What is rail-optimized topology?
Rail-optimized means each GPU server has multiple NICs (typically 8, one per rail) and every rail homes onto its own leaf, so same-index GPUs across servers land on the same leaf and the dominant AllReduce traffic stays local. It is a wiring and traffic-locality discipline layered on a 1:1 non-blocking fabric. It changes where traffic flows, not the leaf and spine switch counts.
How many transceivers does an AI fabric need?
Every point-to-point fabric link uses two transceivers, one at each end, so the transceiver count is the leaf-to-spine link count times two. The link count equals the number of leaves times their uplinks, which also equals the number of spines times their downlinks. This tool reports those quantities. They are planning counts only, not a bill of materials or a quote.
Is RoCEv2 or InfiniBand better for AI?
Neither is universally better. InfiniBand leads on raw performance at low tuning effort and offers a single fabric manager. Ethernet with RoCEv2 matches that performance once tuned and generally wins on cost, openness, multi-tenancy, and industry trajectory. The right fabric depends on cluster size, staffing, and existing stack. This tool weighs the factors for your inputs, then hands the decision to a design review.
How do you make an AI fabric lossless for RoCEv2?
Lossless RoCEv2 is built from three mechanisms working together: PFC for a no-drop class, ECN for early congestion signalling, and ETS for class isolation. DCQCN, which is ECN plus PFC, is the primary steady-state control, and PFC is the last-resort backstop, so tune the ECN marking threshold below the PFC threshold. Keep PFC, ECN, and CoS consistent end to end. OcNOS-DC supports these on the Tomahawk AI platforms.
Does the same OcNOS image run on every switch in the fabric?
Yes. Every leaf and spine switch runs a single OcNOS-DC image with RoCEv2, PFC and ECN, and dynamic load balancing, on open Tomahawk hardware. That keeps the fabric on one operating system and one support contract regardless of how many switches the design produces.