AI data center fabric architecture

An AI data center is not one network. It is four fabrics plus a way to reach across sites, and the whole design lives or dies on how the GPU plane behaves under load. This page puts the pieces together: the AI back-end, storage, front-end, and out-of-band management fabrics, plus coherent DCI, all running on OcNOS-DC over validated Broadcom Tomahawk hardware.

What are the fabrics in an AI data center? An AI data center runs four network planes: the AI back-end fabric that carries GPU-to-GPU collectives, a storage fabric for checkpoints and datasets, a front-end fabric for north-south and tenant access, and an isolated out-of-band management plane. Coherent DCI extends training across sites, and OcNOS-DC runs every one of them on a single NOS.

Design the planes, not just the switches

Most AI fabric mistakes are scoping mistakes: teams size the GPU plane, then bolt storage, tenant access, and management onto it as an afterthought. Separate traffic by plane instead, and give each the subscription ratio and isolation it needs. Only the AI back-end fabric has to be non-blocking and lossless; the other three exist to keep traffic off it that does not belong. All four planes, plus the coherent links that join sites, run on OcNOS-DC on HCL-listed Tomahawk hardware.

The four fabrics

Each plane has one job. Size it for that job, keep its traffic on its own wires, and the fabric stays predictable under a full training run.

Plane 1

AI back-end fabric

The GPU-to-GPU plane that carries collectives. Built 1:1 non-blocking as a rail-optimized leaf-spine pod or a 3-stage Clos, and kept lossless with the RoCEv2 fabric RDMA requires: PFC (including over L3, with DCBX/LLDP) and ECN. DLB spreads flows so no link becomes a hot spot during AllReduce.

1:1 non-blocking · RoCEv2 (PFC + ECN) · DLB · TH5 800G
Plane 2

Storage fabric

A leaf-spine Clos that moves checkpoints and datasets between the GPU racks and NVMe-oF or NFS storage. Typically sized near 3:1, since storage bursts tolerate more oversubscription than collectives, and kept lossless so RDMA storage traffic does not drop under load.

Leaf-spine Clos · ~3:1 · lossless RDMA · NVMe-oF / NFS
Plane 3

Front-end fabric

The north-south plane for tenant access, inference serving, and the path out to the rest of the data center and the internet. Runs 400G or 800G access, sized to the service traffic it carries rather than to the collective pattern of the back-end plane.

North-south · tenant / inference · 400G / 800G
Plane 4

Out-of-band management

A separate Clos or MLAG plane on its own wires, so bring-up and monitoring never share fate with the data planes. Carries ZTP, SNMP and syslog, and gNMI/OpenConfig streaming telemetry, and keeps working when a data-plane fabric is being reconfigured.

Isolated plane · ZTP · SNMP / syslog · gNMI telemetry
Reference architecture

Four fabrics plus DCI, on one NOS

One data center, four planes, extended to a second site over coherent optics. The AI back-end plane is 1:1 non-blocking; storage, front-end, and management each run at the ratio their traffic needs. Every plane runs OcNOS-DC, and the ZR+ DCI link scales training across sites without external transponders.

AI data center fabric architecture: four planes plus coherent DCI to a second site A single AI data center shown as four stacked network planes: an AI back-end fabric (1:1 non-blocking, RoCEv2 lossless, DLB), a storage fabric (roughly 3:1, NVMe-oF and NFS), a front-end fabric (400G and 800G north-south), and an isolated out-of-band management plane (ZTP, SNMP, syslog, gNMI telemetry). All four planes run OcNOS-DC. A 400G/800G ZR+ coherent DCI link on the UfiSpace S9321-64EO connects the site to a second data center, with multi-DC training reach commonly under about 30 km on ZR+ and longer on OpenZR+. Bottom band: one NOS across AI back-end, storage, front-end, out-of-band management, and coherent DCI. DATA CENTER A · OcNOS-DC AI back-end fabric GPU collectives · rail-optimized / Clos 1:1 non-blocking · RoCEv2 · DLB Storage fabric checkpoints / datasets · lossless RDMA ~3:1 · NVMe-oF / NFS Front-end fabric north-south · tenant / inference 400G / 800G Out-of-band mgmt ZTP · SNMP / syslog · telemetry isolated plane · gNMI 数据中心 B AI back-end fabric 1:1 non-blocking · RoCEv2 Storage + front-end ~3:1 · 400G / 800G Out-of-band mgmt ZR+ DCI 400G / 800G ZR+ COHERENT · S9321-64EO · UNDER ~30 KM ZR+ · OPENZR+ FARTHER (oFEC) ONE NOS · OcNOS-DC · AI BACK-END · STORAGE · FRONT-END · OOB · COHERENT DCI

OcNOS 组件: a routed eBGP-unnumbered underlay per plane, RoCEv2 lossless (PFC + ECN, PFC over L3 with DCBX/LLDP) and DLB on the back-end and storage planes, an isolated management plane for ZTP and gNMI telemetry, and 400G/800G ZR+ coherent optics on the S9321-64EO for DCI. Built on HCL-listed Tomahawk 5 (Edgecore AIS800-64D, UfiSpace S9321-64E / 64EO) and Tomahawk 4 (Edgecore AS9736-64D) hardware.

Scale across sites: coherent DCI

When a single training run outgrows one data hall, the fabric extends across the WAN on 400G and 800G ZR+ coherent pluggable optics, no external transponders. The UfiSpace S9321-64EO is the Tomahawk 5 platform that adds 400G ZR+ coherent optics for DCI. Reach is commonly under about 30 km on ZR+, and longer on OpenZR+ using oFEC: close enough to keep collective latency in budget while each site keeps its own leaf-spine fabric unchanged. See the coherent DCI deep-dive for optics and reach detail.

The power of one NOS

The four planes and the DCI links are not four products. They are one operating system in four roles: OcNOS-DC runs the AI back-end, storage, front-end, and out-of-band management fabrics, plus coherent DCI, with one configuration model and one telemetry stack over gNMI and OpenConfig. The team learns one CLI, one automation surface, and one set of counters for every plane.

  • One operational model. The same routing, QoS, and telemetry model applies whether a port is a GPU-facing back-end leaf, a storage leaf, a front-end border, or a management switch.
  • One support contract. A single IP Infusion contract covers the software and the validated hardware, with one TAC and one SLA across every plane. No finger-pointing between a NOS vendor and a hardware vendor.
  • One hardware list. Every plane is built from the same OcNOS 硬件兼容性列表, so support, optics, and firmware are consistent across the design.
  • One roadmap to grow into. On the roadmap: fabric-wide GLB (OcNOS 7.1), latency-based ECN (7.1.0), and LLR, CBFS, and packet trimming. The upcoming Tomahawk 6 silicon (BCM78910 / BCM78914, 102.4 Tbps, TSMC 3nm) extends the same design at higher radix. Build the telemetry plane in from day one so these arrive as software steps.
常见问题

AI fabric architecture FAQ

What fabrics make up an AI data center?
An AI data center runs four network planes. The AI back-end fabric carries GPU-to-GPU collective traffic and is the one that has to be non-blocking and lossless. The storage fabric moves checkpoints and datasets to and from the GPU racks over RDMA. The front-end fabric handles north-south, tenant, and inference access. A separate out-of-band management plane brings switches up and streams telemetry on its own wires. Coherent DCI extends the design across sites. OcNOS-DC runs all four planes plus DCI.
What oversubscription should each fabric use?
The AI back-end fabric should be 1:1 non-blocking on the GPU plane, because a hot link there stalls the whole collective and every GPU waits on the slowest transfer. The storage fabric typically runs a cost-optimized ratio near 3:1, since storage bursts are less latency-critical than collectives. The front-end and out-of-band management planes carry lighter, less bursty traffic and tolerate higher oversubscription. Match the ratio to the traffic each plane actually carries.
How do you connect two AI data centers?
Use coherent DCI: 400G and 800G ZR+ pluggable optics on a border port, no external transponders. On OcNOS-DC the UfiSpace S9321-64EO adds 400G ZR+ coherent optics for exactly this. Multi-DC training reach is commonly under about 30 km on ZR+, and longer on OpenZR+ using oFEC. That lets a single training run span more than one data hall while each site keeps its own leaf-spine fabric unchanged.
Can one NOS run all the fabrics?
Yes. OcNOS-DC runs the AI back-end, storage, front-end, and out-of-band management planes, plus coherent DCI, on the same operating system. That means one configuration model, one telemetry stack over gNMI and OpenConfig, and one IP Infusion support contract covering the software and the validated hardware, with one TAC and one SLA across every plane in the design.
What hardware runs each plane?
All four planes run OcNOS-DC on Broadcom Tomahawk silicon. The 800G planes use Tomahawk 5 (BCM78900, 51.2 Tbps, 64×800G): Edgecore AIS800-64D or UfiSpace S9321-64E, with the S9321-64EO adding 400G ZR+ coherent optics for DCI. Entry and 400G planes use Tomahawk 4 (BCM56990, 25.6 Tbps, 64×400G) such as the Edgecore AS9736-64D. These are on-chip shared-buffer switches, and every one is listed on the OcNOS Hardware Compatibility List.

Designing the full AI data center? We'll plan every plane with you.

预约架构评审 →