AI fabric design tool
Design a non-blocking RoCEv2 GPU fabric end to end, vendor-neutral. Size the leaf and spine switches, plan the transceiver and cable counts, compare InfiniBand against Ethernet for your cluster, and get a lossless congestion profile. This designs the network only. It is not a cost estimate or a bill of materials.
Size the GPU fabric
Pick a starting size, or enter your own GPU count, port speed, and switch silicon. The tool sizes a folded two-tier leaf-spine pod and flags when a cluster crosses into three-tier scale.
…
Counts assume a folded leaf-spine (Clos) fabric with one fabric NIC per GPU, where each leaf splits its ports between GPUs and spine uplinks. Rail-optimized wiring (multiple NICs per GPU) changes traffic locality, not these switch counts. See the AI fabric topology reference designs for the rail-optimized and three-tier layouts.
Component summary
Fabric link, transceiver, and cable quantities for the sized design. These are counts for planning only, not a bill of materials, quote, or pricing.
| Component | Basis | Quantity |
|---|---|---|
| GPU-facing links | GPUs × 1 NIC | … |
| Leaf to spine fabric links | leaves × uplinks | … |
| Fabric transceivers | links × 2 ends | … |
| Fabric cables | 1 per link | … |
Media by reach, chosen per your cabling: up to 3 m DAC or AEC, up to 50 to 100 m AOC or SR, up to 500 m DR (the workhorse), up to 2 km FR. Reach classes only. No pricing, SKUs, or vendor part numbers. On a three-tier design, add the spine to super-spine links on top of the counts above.
InfiniBand vs Ethernet
A vendor-neutral scorecard across six factors. It shows where each factor leans for your inputs and why. There is no single winner here by design. Use it as an input to a design review, not a final recommendation.
| Factor | InfiniBand | Ethernet / RoCEv2 | Ultra Ethernet (trajectory) |
|---|---|---|---|
| Peak performanceleans InfiniBand | A mid-teens percent raw throughput edge at low tuning effort, plus SHARP in-network reduction that offloads collective operations from the GPUs. | Matches InfiniBand with proper tuning. A 24,000-GPU RoCEv2 training cluster has run in production at performance parity. The gap is tuning effort, not the ceiling. | Adds packet spray and NIC-side reordering to cut tail latency and reduce PFC dependence, narrowing the out-of-box gap. |
| Costleans Ethernet | Carries a meaningful premium at fabric scale, commonly cited in the 30 to 60 percent range versus comparable Ethernet. | Merchant-silicon switches and a broad optics supply chain generally make it the lower-cost fabric per port. | Same open merchant-silicon economics as Ethernet. |
| Opennessleans Ethernet | Single-vendor stack end to end (switches, NICs, manager), so multi-vendor sourcing is limited. | Open and multi-vendor: standards-based switches, NICs, and NOS from many suppliers, preserving supply-chain optionality. | Industry consortium standard (UEC 1.0, 2025) reinforcing the open direction. |
| Operational simplicityneutral | UFM gives a single, purpose-built fabric manager, which is turnkey for teams without deep RDMA staff. | Rides the large Ethernet talent pool and standard tooling, but lossless RoCEv2 needs deliberate end-to-end tuning. | Aims to reduce the tuning burden, moving Ethernet toward IB-like out-of-box behavior. |
| Multi-tenancy and convergenceleans Ethernet | Typically a dedicated back-end island, without native multi-tenant isolation or storage convergence. | One converged Ethernet fabric can carry back-end, storage, and front-end traffic with standard tenant isolation. | Extends the converged-Ethernet direction on one open fabric. |
| Trajectoryleans Ethernet | Mature and stable, performance-leading at low tuning, but on a single-vendor roadmap. | Ethernet crossed over to lead AI back-end deployments by mid-2025, with broad ecosystem momentum. | UEC 1.0 (2025) is a multi-vendor spec explicitly targeting AI and HPC back-end fabrics. |
Assumptions: performance is compared with tuning (untuned Ethernet trails, well-tuned Ethernet reaches parity); cost is a range, never a dollar figure; openness means multi-vendor sourcing, not a quality judgment; Ultra Ethernet reflects a 2025-onward direction, not shipping product. Collective tuning refers to xCCL libraries generically. This informs the decision; validate a fabric choice with an IP Infusion engineer.
RoCEv2 lossless profile
RoCEv2 is not a single switch feature. It is a profile composed from PFC, ECN, and ETS, with DCQCN doing the steady-state work and PFC as a backstop. This recommends the mechanisms and the order to apply them for your design. It does not emit threshold values and it does not touch any device.
| Mechanism | Role in the lossless profile | OcNOS availability |
|---|---|---|
| PFC | No-drop guarantee for the RoCE class. Last-resort backstop, not the primary control. | All 4 AI platforms |
| ECN | Marks congestion early so senders throttle before queues overflow. Tune its marking point below the PFC trigger. | All 4 AI platforms |
| DCQCN (ECN + PFC) | Primary steady-state control. ECN marking drives the DCQCN reaction at the NICs; PFC backs it up. Composed, not a standalone feature. | All 4 AI platforms |
| ETS | Bandwidth guarantees and class isolation so the lossless class is not starved by best-effort traffic. | All 4 AI platforms |
| PFC-deadlock watchdog | Detects and breaks PFC pause deadlock or storm conditions, the safety net for the PFC backstop. | All 4, OcNOS 7.0 |
| DLB (Dynamic Load Balancing) | Spreads elephant RoCE flows that collide under static ECMP. Lifts fabric utilization from roughly 55 percent toward 90 percent and above. | All 4 AI platforms |
| Dynamic ECN | Adapts ECN marking to live queue conditions, reducing manual re-tuning as load shifts. | TH5 only, OcNOS 7.0 |
| DLB reactive / random | Advanced DLB placement modes for tighter flow spreading. On TH4, use standard DLB. | TH5 only, OcNOS 7.0 |
| GLB (Global Load Balancing) | Fabric-wide load balancing beyond local DLB decisions. | Roadmap, OcNOS 7.1 train |
| Ultra Ethernet | Packet spray with NIC-side reordering and reduced PFC dependence, the trajectory for lossless AI Ethernet. | Roadmap, UEC profile |
Advisory only. These recommendations are not pushed to any device and no threshold values are generated here. Availability is per the OcNOS-DC Feature Matrix and Hardware Compatibility List. Validate against your selected platforms and OcNOS release before deployment.
Fabric power
A typical power range for the fabric switches in the sized design. Ranges only, optics-loaded. Network switches are a small part of total cluster power; the GPUs dominate.
Cooling: rear-door or direct-liquid cooling is typically considered above roughly 30 to 40 kW per rack, driven by GPU-server density, not the network. This is guidance, not a computed thermal load. Switch power ranges are typical values for 51.2T (Tomahawk 5) and 25.6T (Tomahawk 4) class boxes with optics populated; confirm exact figures on each platform datasheet.
Matching OcNOS-DC platforms
The fabric runs a single OcNOS-DC image on open Broadcom Tomahawk hardware. These are the supported 800G and 400G platforms it designs onto.




This tool provides a structural network-design estimate for planning only. It is not a performance guarantee, a bill of materials, or a cost estimate. Actual designs depend on rail optimization, NIC count per GPU, cabling, and failure-domain choices. GPU counts are reference-design ceilings from Clos radix math, not measured figures. Broadcom and Tomahawk are trademarks of Broadcom Inc.; other names are the trademarks of their respective owners.
AI Fabric Reference Design
Get the reference design PDF: topology, switch counts, component summary, and a RoCEv2 lossless starting profile.
AI Fabric Reference Design
Quick form: your PDF will download immediately after submit.
✓ Opening your PDF in a new tab…
If it didn't open, use the link below.
AI Fabric Reference Design (PDF)