AI-Fabric-Topologien: Rail-Optimized- und Scheduled-Designs

The shape of your fabric decides the shape of your training job. This page lays out the reference topologies OcNOS-DC ships against, from a rail-optimized single pod, through a scheduled 3-stage Clos, to coherent multi-DC DCI, sized in concrete port-counts on Broadcom Tomahawk 4 and Tomahawk 5 hardware.

Which AI fabric topology should you use? Pick the smallest non-blocking design that keeps every GPU's link saturated during collectives. Up to about 1,000 GPUs, use a rail-optimized leaf-spine pod (8 rails per GPU server, one rail per leaf). From roughly 1,000 to 16,000+ GPUs, move to a 3-stage Clos (leaf, spine, super-spine). To span data centers, extend with 400G ZR/ZR+ coherent DCI. All three run on OcNOS-DC over a RoCEv2 lossless (PFC + ECN) L3 fabric.

Wählen Sie nach GPU-Anzahl, nicht nach Buzzword

Eine AI-Fabric-Topologie hat eine Aufgabe: halten every Der Outbound-Link der GPU wird während einer Collective-Operation ausgelastet, ohne Ausreißer bei der Tail-Latenz zu erzeugen. Die richtige Topologie ist die kleinste, die dies für Ihre GPU-Anzahl leistet, mit einem Fallback-Pfad für die nächstgrößere Ausbaustufe. Nachfolgend: drei Referenzdesigns, die OcNOS-DC heute unterstützt, mit konkreten Port-Berechnungen.

Dimensionieren Sie Ihr eigenes Cluster? The AI Fabric Design Suite liefert eine schnelle erste Einschätzung: Es dimensioniert einen nichtblockierend, zweistufig Leaf-Spine-Pod unter der Annahme ein Fabric-NIC pro GPU, und weist darauf hin, wenn Sie in den dreistufigen Maßstab übergehen. Die nachstehenden Referenzdesigns verwenden das same die nicht blockierende Leaf/Spine-Berechnung und deren Erweiterung auf ein 3-stufiges Clos im großen Maßstab, sodass die Switch-Anzahl mit dem Werkzeug übereinstimmt. Rail-optimized here is the wiring discipline of an 8-NIC GPU server (one rail per leaf, so intra-rail AllReduce stays on the leaf) layered on that non-blocking fabric: it changes traffic locality, not the switch count. Use the tool for a ballpark; use these designs for the build.

256GPUs

Einstiegs-Non-Blocking-Pod

Eine Rack-Reihe rail-aligned Leaves über einer kleinen Spine-Ebene. Zwei-Stufen-Folded-Clos, 1:1 non-blocking.

8 Leaves · 4 Spines · TH4 · 400G
1,024GPUs

Rail-optimierter 2-Tier-Pod

Rail-ausgerichtete Leaves mit einem 1:1 non-blocking Spine. Intra-Rail-AllReduce bleibt auf dem Leaf, Cross-Rail-Traffic nutzt den Spine. Die standardmäßige skalierbare Single-Pod-Einheit.

32 Leaves · 16 Spines · TH5 · 800G
4,096GPUs

3-stufiges Clos

Leaf, spine, super-spine. Each 1,024-GPU pod is 1:1 non-blocking; a super-spine plane scales across pods. DLB at every tier; GLB end-to-end on the OcNOS 7.1 train.

128 leaves · 64 spines · 32 super-spines · TH5 · 800G
16,384GPUs

Skaliertes 3-stufiges Clos

Multi-Pod-3-Stage-Clos mit Super-Spine-Plane. Dimensioniert für die Training-Klasse mit Billionen Parametern.

512 leaves · 256 spines · 128 super-spines · TH5 · 800G
Referenzdesign 1

Rail-optimierter Single Pod

Jeder GPU-Server hat 8 NICs, eine pro „Rail“ (ein dedizierter xCCL-Collective-Channel (NCCL / RCCL / oneCCL)). Jede Rail hat ihr eigenes dediziertes Leaf – damit landen alle 8 NICs jedes Servers auf einem unterschiedlichen Leaf. AllReduce über Rail-N bleibt innerhalb von Leaf-N. Keinerlei East-West-Druck auf den Spine für das dominierende Collective-Pattern.

Rail-optimiertes AI-Fabric: 8 Rails, rail-aligned Leaves, 1:1 non-blocking Spine (schematisch) Rail-optimierte AI-Fabric. Acht GPU-Server am unteren Rand verfügen jeweils über acht NICs, die auf acht Rail-Leaves ausgerichtet sind. Rail-N jedes Servers verbindet sich mit Leaf-N. Eine Spine-Ebene oberhalb der Leaves transportiert den Cross-Rail-Traffic. Der dominante AllReduce-Traffic bleibt innerhalb einer Rail und durchquert niemals den Spine. Spine-1TH5 · 800G Spine-2TH5 · 800G Spine-3TH5 · 800G Spine-4TH5 · 800G Rail-1leaf Rail-2leaf Rail-3leaf Rail-4leaf Rail-5leaf Rail-6leaf Rail-7leaf Rail-8leaf GPU-Server 1 8 × NIC · 8 Rails GPU-Server 2 8 × NIC · 8 Rails GPU-Server 3 8 × NIC · 8 Rails GPU-Server 4 8 × NIC · 8 Rails RAIL-OPTIMIZED · 8 RAILS · INTRA-RAIL ALLREDUCE STAYS LOCAL

OcNOS-Komponenten: BGP-unnumbered L3 underlay, RoCEv2 lossless (PFC + ECN) on every leaf, DLB at the spine tier. Built on HCL-listed hardware: the 800G scalable unit uses TH5 64×800G leaves and spines (Edgecore AIS800-64D or UfiSpace S9321-64E); the entry 256-GPU pod uses Edgecore AS9736-64D (TH4, 64×400G).

Scheduled vs. Rail-Aligned: was sich im großen Maßstab ändert

Rail-optimized stops scaling somewhere between 1k and 2k GPUs: you run out of leaf radix, or the spine tier becomes too oversubscribed. Above that, most modern AI fabrics move to a 3-stage Clos: leaf, spine, super-spine. The hard part on a Clos is spreading flows evenly so no link becomes a hot spot. Approaches run from per-flow ECMP, through adaptive (dynamic) load balancing, to per-packet spray, the model Ultra Ethernet uses with reordering handled at the NIC. A separate family, cell-based scheduled fabrics such as Broadcom DDC, segments traffic into cells and schedules it inside the fabric. OcNOS keeps the GPU plane balanced with DLB today and adds fabric-wide GLB on the 7.1 train, and is UEC-ready as UEC NICs arrive.

Referenzdesign 2

3-Stage Clos Scheduled Fabric: 4,096-16,384 GPUs

Three tiers: leaf, spine, super-spine. Any two GPUs are at most four switch hops apart; same-leaf and same-pod peers are closer. Non-blocking within a pod, with the super-spine plane setting the cross-pod ratio. DLB at every hop, GLB across the full path on the OcNOS 7.1 train, UEC packet-spray on UEC-capable NICs. The diagram is schematic: it draws a reduced tier count; the 4,096-GPU build is 128 leaves / 64 spines / 32 super-spines on TH5 800G.

3-stufige Clos-AI-Fabric mit Scheduled Topology Dreistufige Clos-Topologie. Die obere Ebene zeigt vier Super-Spine-Switches. Die mittlere Ebene zeigt acht Spine-Switches. Die untere Ebene zeigt 12 Leaf-Switches, die GPU-Pods speisen. Vollständige Mesh-Verbindungen von Leaf zu Spine und von Spine zu Super-Spine. Beschriftungen des unteren Bandes: 4096-GPU-Scheduled-Fabric, DLB auf jeder Ebene, GLB Ende-zu-Ende mit OcNOS 7.1. Super-Spine-1 Super-Spine-2 Super-Spine-3 Super-Spine-4 Spine-1 Spine-2 Spine-3 Spine-4 Spine-5 Spine-6 Spine-7 Spine-8 L1 L2 L3 L4 L5 L6 L7 L8 L9 L10 L11 L12 SUPER-SPINE SPINE LEAF GPU PODS 128 Leaves · 32 GPUs/Leaf · 4.096 GPUs gesamt · TH5 · 800G 3-STAGE CLOS · 4.096 GPU · DLB EVERY HOP · GLB E2E (OcNOS 7.1) · UEC-READY

OcNOS-Komponenten: eBGP-Unnumbered-L3-Underlay, verlustfreies RoCEv2 (PFC + ECN), DLB auf jeder Ebene, GLB Ende-zu-Ende auf dem OcNOS-7.1-Train, gNMI-Streaming-Telemetry an Ihren Observability-Stack; EVPN-VXLAN-Multi-Tenant-Overlay, wo Mandantenfähigkeit erforderlich ist. Durchgehend auf HCL-gelisteten TH5-64×800G-Chassis aufgebaut.

Subscription is a dial, not a fixed rule. These counts make each 1,024-GPU pod 1:1 non-blocking and use a cost-optimized ~2:1 super-spine for cross-pod traffic, the rail-optimized approach hyperscale Ethernet fabrics rely on (published large-scale designs oversubscribe the top tier far more, because collective traffic stays pod-local). Want maximal any-to-any headroom instead? A fully non-blocking 1:1 build is 128 / 128 / 64 at 4,096 GPUs and 512 / 512 / 256 at 16,384; only the spine and super-spine counts change. Model either in the AI Fabric Design Suite.

Multi-DC und DCI für verteiltes Training

Wenn ein einzelner Trainingslauf über mehr als eine Data Hall reicht, was bei Modellen mit einer Billion Parametern zunehmend üblich wird, erstreckt sich die Fabric über das WAN. OcNOS-DC unterstützt 400G ZR / ZR+ Coherent-Optiken direkt auf dem Spine für transponderfreie DCI, wobei die EVPN-Tunnelerweiterung VXLAN-Tenants über Standorte hinweg transportiert.

Referenzdesign 3

Multi-DC-AI-Fabric: Coherent DCI

Zwei AI-Rechenzentren, am Spine über 400G ZR/ZR+ zusammengeführt. EVPN inter-DC trägt die L2/L3-Mandantenerweiterung; der zugrunde liegende 3-stufige Clos in jedem Standort bleibt unverändert.

Multi-DC-AI-Fabric mit 400G ZR/ZR+ DCI Zwei AI-Rechenzentren, jedes mit einer Leaf-Spine-Fabric. Die beiden Spines sind über kohärente 400G ZR/ZR+ Optik über ein WAN verbunden. EVPN-Inter-DC-Tunnel dehnen Tenants von einem Standort zum anderen aus. Unteres Band: transponderfreies kohärentes DCI. DATA CENTER A DATA CENTER B Spine-A1400G ZR+ Spine-A2400G ZR+ Spine-B1400G ZR+ Spine-B2400G ZR+ EVPN inter-DC · 400G ZR/ZR+ Leaf-A1 Leaf-A2 Leaf-A3 Leaf-B1 Leaf-B2 Leaf-B3 GPU-Pods · Site A GPU-Pods · Site B COHERENT DCI · TRANSPONDER-FREI · EVPN INTER-DC · 400G ZR/ZR+

OcNOS-Komponenten: 400G ZR/ZR+ pluggable coherent optics on a DWDM-capable spine or border-leaf port, EVPN inter-DC for tenant L2/L3 extension, gNMI telemetry across sites. No external transponders required. Reach: 400ZR to roughly 120 km amplified; OpenZR+ reaches farther on oFEC.

Design-Faustregeln

  • Passen Sie die Topologie an die GPU-Anzahl an. Kleinste Pods (unterhalb der NIC-Radix eines Leaf): Rail-only genügt. Single-Pod-Maßstab: Rail-optimiertes Leaf-Spine. Multi-Pod: Ein dreistufiges Clos ist das einzige Design, das ohne Oversubscription-Kompromisse skaliert.
  • Immer 1:1-Subscription auf der AI-Plane. Storage- und CPU-Racks können höhere Oversubscription-Ratios fahren. Die GPU-Plane sollte das nicht.
  • Planen Sie die Rail-Anzahl von xCCL aus, nicht nach Verkabelungs-Komfort. 8 Rails ist der derzeitige De-facto-Standard für 8-NIC-GPU-Server. Kombinieren Sie Rails nicht zu weniger Leaves.
  • Wählen Sie das Silicon nach Leistung und Dichte, nicht nach dem Label. TH4 (25,6T) und TH5 (51,2T) sind die Arbeitspferde; die Wahl zwischen beiden hängt von der Rack-Leistung und den Kosten für Breakout-Kabel ab.
  • GLB / UEC bereits in der Designphase einplanen. Bauen Sie die Telemetry-Plane von Tag eins an mit ein, auch auf einer 7.0-Fabric, sodass das OcNOS-7.1-GLB-Upgrade rein ein Software-Schritt ist. Siehe GLB und Ultra Ethernet.
  • Gegen die HCL validieren. Jede Referenz hier basiert auf Hardware, die in der OcNOS Hardware Compatibility List; wählen Sie von dort aus für erstklassigen Support.
FAQ

AI fabric topology FAQ

What is a rail-optimized topology, and how is it different from rail-only?
Rail-optimized wiring connects each of a GPU server's 8 NICs to its own dedicated rail leaf, so the dominant same-rail AllReduce traffic stays on one leaf and never traverses the spine. Rail-only is the small-cluster case: a single rack-row of rail-aligned leaves with no spine tier, where cross-rail traffic relies on the GPU scale-up domain. Rail-optimized adds a non-blocking spine so cross-rail flows have a network path.
How many GPUs can a 3-stage Clos scale to?
On radix-64 switches (Tomahawk 5 at 800G) a 2-tier leaf-spine tops out at 2,048 GPUs. A 3-stage Clos with a super-spine tier extends that to 16,000+ GPUs in the reference designs above, and up to about 65,000 GPUs at the theoretical fat-tree limit. Because most collective traffic stays rail-local, the super-spine plane is sized to the cross-pod ratio you actually need.
Should I use Tomahawk 4 or Tomahawk 5?
Both run OcNOS-DC. Tomahawk 4 (25.6 Tbps, 64×400G) is the cost-optimized choice for entry pods and 400G GPU NICs. Tomahawk 5 (51.2 Tbps, 64×800G) is the workhorse for 800G GPU servers and larger fabrics. Tomahawk 4 has no native 800G, so match the switch to your NIC speed.
Do I need InfiniBand, or is Ethernet enough?
Ethernet is now a first-class AI-fabric transport. RoCEv2 with PFC and ECN delivers lossless RDMA today. Ultra Ethernet (UEC) removes the network-wide PFC dependency using endpoint packet-spray, selective retransmission, and link-level retry as UEC NICs ship. OcNOS-DC runs the RoCEv2 fabric today and is UEC-ready.
Where does scale-up end and scale-out begin?
Inside a GPU server and its NVLink domain (for example GB200 NVL72), GPUs communicate over the scale-up fabric at terabit speeds. The rail, leaf-spine, and Clos network is the scale-out fabric between servers and pods. Most same-rail collective traffic is absorbed by scale-up first, so the network carries the cross-rail and cross-pod remainder, which is why 1:1 non-blocking matters most on the GPU plane.

Sie konzipieren Ihr AI-Fabric? Wir berechnen die Port-Anzahl gemeinsam mit Ihnen.

Architecture Review buchen →
AI-Fabric

Design the whole AI fabric with OcNOS

From the business case to the port-count maths, pick up wherever you are in the build.