AI back-end · Storage · Front-end · Coherent DCI

AI data center fabric architecture

An AI data center runs four network planes, not one flat network: an AI back-end fabric for GPU-to-GPU collectives, a storage fabric, a front-end fabric, and an isolated out-of-band management plane. Coherent DCI extends training across sites, and OcNOS-DC runs every plane on a single NOS over validated Broadcom Tomahawk hardware.

4 planesplus coherent DCI
1 NOSOcNOS-DC on every plane
up to 800GTomahawk 5, 64x800G
1 contractone TAC, one SLA
Scope by plane

Design the planes, not just the switches

Most AI fabric mistakes are scoping mistakes: teams size the GPU plane, then bolt storage, tenant access, and management onto it as an afterthought.

Separate traffic by plane instead, and give each the subscription ratio and isolation it needs. Only the AI back-end fabric has to be non-blocking and lossless; the other three exist to keep traffic off it that does not belong. All four planes, plus the coherent links that join sites, run on OcNOS-DC on HCL-listed Tomahawk hardware.

The four fabrics

Each plane has one job

Size each plane for its own job, keep its traffic on its own wires, and the fabric stays predictable under a full training run.

Plane 1

AI back-end fabric

The GPU-to-GPU plane that carries collectives. Built 1:1 non-blocking as a rail-optimized leaf-spine pod or a 3-stage Clos on a pure Layer 3 eBGP underlay, and kept lossless with the RoCEv2 fabric RDMA requires: PFC (including over L3, with DCBX/LLDP) and ECN. DLB spreads flows so no link becomes a hot spot during AllReduce.

1:1 non-blocking · RoCEv2 (PFC + ECN) · DLB · TH5 800G
Plane 2

Storage fabric

A leaf-spine Clos that moves checkpoints and datasets between the GPU racks and NVMe-oF or NFS storage. Typically sized near 3:1, since storage bursts tolerate more oversubscription than collectives, and kept lossless so RDMA storage traffic does not drop under load.

Leaf-spine Clos · ~3:1 · lossless RDMA · NVMe-oF / NFS
Plane 3

Front-end fabric

The north-south plane for tenant access, inference serving, and the path out to the rest of the data center and the internet. Runs 400G or 800G access, sized to the service traffic it carries rather than to the collective pattern of the back-end plane.

North-south · tenant / inference · 400G / 800G
Plane 4

Out-of-band management

A separate Clos or MLAG plane on its own wires, so bring-up and monitoring never share fate with the data planes. Carries ZTP, SNMP and syslog, and gNMI/OpenConfig streaming telemetry, and keeps working when a data-plane fabric is being reconfigured.

Isolated plane · ZTP · SNMP / syslog · gNMI telemetry
Reference architecture

Four fabrics plus DCI, on one NOS

One data center, four planes, extended to a second site over coherent optics. The AI back-end plane is 1:1 non-blocking; storage, front-end, and management each run at the ratio their traffic needs. Every plane runs OcNOS-DC, and the ZR+ DCI link scales training across sites without external transponders.

AI data center fabric architecture: four planes plus coherent DCI to a second site A single AI data center shown as four stacked network planes: an AI back-end fabric (1:1 non-blocking, RoCEv2 lossless, DLB), a storage fabric (roughly 3:1, NVMe-oF and NFS), a front-end fabric (400G and 800G north-south), and an isolated out-of-band management plane (ZTP, SNMP, syslog, gNMI telemetry). All four planes run OcNOS-DC. A 400G/800G ZR+ coherent DCI link on the UfiSpace S9321-64EO connects the site to a second data center, with multi-DC training reach commonly under about 30 km on ZR+ and longer on OpenZR+. Bottom band: one NOS across AI back-end, storage, front-end, out-of-band management, and coherent DCI. DATA CENTER A · OcNOS-DC AI back-end fabric GPU collectives · rail-optimized / Clos 1:1 non-blocking · RoCEv2 · DLB Storage fabric checkpoints / datasets · lossless RDMA ~3:1 · NVMe-oF / NFS Front-end fabric north-south · tenant / inference 400G / 800G Out-of-band mgmt ZTP · SNMP / syslog · telemetry isolated plane · gNMI DATA CENTER B AI back-end fabric 1:1 non-blocking · RoCEv2 Storage + front-end ~3:1 · 400G / 800G Out-of-band mgmt ZR+ DCI 400G / 800G ZR+ COHERENT · S9321-64EO · UNDER ~30 KM ZR+ · OPENZR+ FARTHER (oFEC) ONE NOS · OcNOS-DC · AI BACK-END · STORAGE · FRONT-END · OOB · COHERENT DCI

OcNOS pieces: a routed eBGP-unnumbered underlay per plane (the AI back-end is pure Layer 3 eBGP, with EVPN-VXLAN kept as a tenant overlay rather than the AI underlay), RoCEv2 lossless (PFC + ECN, PFC over L3 with DCBX/LLDP) and DLB on the back-end and storage planes, an isolated management plane for ZTP and gNMI telemetry, and 400G/800G ZR+ coherent optics on the S9321-64EO for DCI. Built on HCL-listed Tomahawk 5 (Edgecore AIS800-64D, UfiSpace S9321-64E / 64EO) and Tomahawk 4 (Edgecore AS9736-64D) hardware.

Scale across sites

Coherent DCI, no external transponders

When a single training run outgrows one data hall, the fabric extends across the WAN on 400G and 800G ZR+ coherent pluggable optics, straight off a border port.

The UfiSpace S9321-64EO is the Tomahawk 5 platform that adds 400G ZR+ coherent optics for DCI. Reach is commonly under about 30 km on ZR+, and longer on OpenZR+ using oFEC: close enough to keep collective latency in budget while each site keeps its own leaf-spine fabric unchanged. See the coherent DCI deep-dive for optics and reach detail.

The silicon

What hardware runs each plane

Every plane runs OcNOS-DC on Broadcom Tomahawk silicon. The 800G planes and DCI ride Tomahawk 5; entry and 400G planes ride Tomahawk 4. Both are on-chip shared-buffer switches, and every platform is on the OcNOS Hardware Compatibility List.

Silicon layer Tomahawk 5800G planes + DCI Tomahawk 4entry + 400G planes
Broadcom partBCM78900BCM56990
Switch capacity51.2 Tbps25.6 Tbps
Port configuration64×800G64×400G
BufferOn-chip shared buffer; HCL-listed.On-chip shared buffer; HCL-listed.
PlatformsEdgecore AIS800-64D, UfiSpace S9321-64E. The S9321-64EO adds 400G ZR+ coherent optics for DCI.Edgecore AS9736-64D.
Role in the designAI back-end and storage planes at 800G, plus coherent DCI across sites.Entry builds and 400G front-end or management planes.
One NOS, four roles

The power of one NOS

The four planes and the DCI links are not four products. They are one operating system in four roles: OcNOS-DC runs the AI back-end, storage, front-end, and out-of-band management fabrics, plus coherent DCI, with one configuration model and one telemetry stack over gNMI and OpenConfig. The team learns one CLI, one automation surface, and one set of counters for every plane.

One operational model

The same routing, QoS, and telemetry model applies whether a port is a GPU-facing back-end leaf, a storage leaf, a front-end border, or a management switch.

One support contract

A single IP Infusion contract covers the software and the validated hardware, with one TAC and one SLA across every plane. No finger-pointing between a NOS vendor and a hardware vendor.

One hardware list

Every plane is built from the same OcNOS Hardware Compatibility List, so support, optics, and firmware stay consistent across the design.

One roadmap to grow into

On the roadmap: fabric-wide GLB (OcNOS 7.1), latency-based ECN (7.1.0), and LLR, CBFS, and packet trimming. Upcoming Tomahawk 6 silicon (BCM78910 / BCM78914, 102.4 Tbps, TSMC 3nm) rides the OcNOS 7.2 train and extends the same design at higher radix. Build the telemetry plane in from day one so these arrive as software steps.

FAQ

AI fabric architecture FAQ

What fabrics make up an AI data center?
An AI data center runs four network planes. The AI back-end fabric carries GPU-to-GPU collective traffic and is the one that has to be non-blocking and lossless. The storage fabric moves checkpoints and datasets to and from the GPU racks over RDMA. The front-end fabric handles north-south, tenant, and inference access. A separate out-of-band management plane brings switches up and streams telemetry on its own wires. Coherent DCI extends the design across sites. OcNOS-DC runs all four planes plus DCI.
What oversubscription should each fabric use?
The AI back-end fabric should be 1:1 non-blocking on the GPU plane, because a hot link there stalls the whole collective and every GPU waits on the slowest transfer. The storage fabric typically runs a cost-optimized ratio near 3:1, since storage bursts are less latency-critical than collectives. The front-end and out-of-band management planes carry lighter, less bursty traffic and tolerate higher oversubscription. Match the ratio to the traffic each plane actually carries.
How do you connect two AI data centers?
Use coherent DCI: 400G and 800G ZR+ pluggable optics on a border port, no external transponders. On OcNOS-DC the UfiSpace S9321-64EO adds 400G ZR+ coherent optics for exactly this. Multi-DC training reach is commonly under about 30 km on ZR+, and longer on OpenZR+ using oFEC. That lets a single training run span more than one data hall while each site keeps its own leaf-spine fabric unchanged.
Can one NOS run all the fabrics?
Yes. OcNOS-DC runs the AI back-end, storage, front-end, and out-of-band management planes, plus coherent DCI, on the same operating system. That means one configuration model, one telemetry stack over gNMI and OpenConfig, and one IP Infusion support contract covering the software and the validated hardware, with one TAC and one SLA across every plane in the design.
What hardware runs each plane?
All four planes run OcNOS-DC on Broadcom Tomahawk silicon. The 800G planes use Tomahawk 5 (BCM78900, 51.2 Tbps, 64×800G): Edgecore AIS800-64D or UfiSpace S9321-64E, with the S9321-64EO adding 400G ZR+ coherent optics for DCI. Entry and 400G planes use Tomahawk 4 (BCM56990, 25.6 Tbps, 64×400G) such as the Edgecore AS9736-64D. These are on-chip shared-buffer switches, and every one is listed on the OcNOS Hardware Compatibility List.

Designing the full AI data center? We will plan every plane with you.

Tell us the GPU scale and the workload, and an IP Infusion engineer will size the back-end, storage, front-end, management, and DCI planes with you, or start with a first-pass layout in the AI Fabric Design Suite.