What Is a GPU Fabric?
A GPU fabric is the back-end network that interconnects GPUs across servers and pods for training and inference. It is the scale-out network beyond the in-server scale-up domain, and it carries the RDMA traffic GPUs exchange.
Scale-up ends, scale-out begins
A GPU fabric sits on one side of a clear boundary.
Scale-up domain
The very high bandwidth interconnect inside a GPU server and its tightly coupled domain, for example NVLink within an NVL72 rack. GPUs here talk at terabit speeds as if they were one large accelerator. This is not the GPU fabric.
Scale-out GPU fabric
The Ethernet network between servers and pods: rail, leaf, spine, and Clos tiers. This is the GPU fabric. A collective uses scale-up first inside the domain, then this fabric carries the cross-server remainder.
Getting this boundary right matters for sizing. Because same-rail traffic is often absorbed by scale-up first, the GPU fabric carries the cross-rail and cross-pod remainder, which is why 1:1 non-blocking is worth paying for on the GPU plane specifically and not on storage or CPU racks.
How the fabric is shaped: rails and Clos
Most GPU fabrics are rail-optimized leaf-spine networks: each of a GPU server's 8 NICs connects to its own dedicated rail leaf, so the dominant same-rail AllReduce traffic stays on one leaf and never crosses the spine. Past a single pod, the design extends to a 3-stufiges Clos (leaf, spine, super-spine). On radix-64 Tomahawk 5 switches, a 2-tier fabric reaches roughly 2,048 GPUs at 1:1 non-blocking and a 3-stage Clos scales to 16,000+ GPUs, up to about 65,535 at the fat-tree limit. See rail-optimized network und AI-Fabric-Topologien for the port maths.
Lossless RoCEv2 on the GPU plane
GPUs move data with RDMA over RoCEv2, which performs poorly under packet loss. So the GPU fabric is engineered lossless with Priority Flow Control and ECN, and kept evenly loaded with dynamic load balancing so no link becomes a hot spot during a collective. OcNOS-DC supplies PFC (including over L3), ECN, Dynamic ECN, and DLB on Broadcom Tomahawk 5 (51.2 Tbps, 64x800G) platforms such as the Edgecore AIS800-64D and UfiSpace S9321-64E. That no-drop, evenly-loaded behavior is what keeps a training step from stalling.
Standing up a GPU fabric? Let's size the plane together.
Architecture Review buchen →What Is a GPU Fabric FAQ
Was ist ein GPU-Fabric?
What is the difference between scale-up and scale-out?
What topology does a GPU fabric use?
Does a GPU fabric have to be lossless?
Design the whole AI fabric with OcNOS
From the business case to the port-count maths, pick up wherever you are in the build.
Gehen Sie tiefer. Nehmen Sie es mit.
Two short, technical downloads that go further than this page: the full OcNOS-DC datasheet and the lossless 800G AI fabric architecture.
OcNOS-DC Datenblatt
Vollständige OcNOS-DC Spezifikation: Funktionsumfang für EVPN-VXLAN und Ethernet for AI, Software-SKUs, unterstützte Hardware-Plattformen und der Solution Ordering Guide.
Datasheet abrufenOcNOS 800G verlustfreie AI Fabric
Non-blocking RoCEv2-Fabric auf Broadcom Tomahawk 4/5 Spines: SKU-Stufen, validierte Plattformen und Deployment-Architektur.
Brief anfordernOcNOS-DC Datenblatt
Kurzes Formular. Ihr PDF öffnet sich unmittelbar nach dem Absenden in einem neuen Tab.
✓ Ihr PDF wird in einem neuen Tab geöffnet…
Falls es sich nicht geöffnet hat, nutzen Sie den untenstehenden Link.
OcNOS 800G verlustfreie AI Fabric
Kurzes Formular. Ihr PDF öffnet sich unmittelbar nach dem Absenden in einem neuen Tab.
✓ Ihr PDF wird in einem neuen Tab geöffnet…
Falls es sich nicht geöffnet hat, nutzen Sie den untenstehenden Link.