Networking Definition · 800G · RoCEv2

What Is an AI Networking Switch?

An AI networking switch is a high-radix Ethernet data center switch, usually 400G or 800G per port, that connects GPU servers in the AI back-end fabric and keeps RDMA traffic lossless with PFC, ECN and adaptive load balancing, because the slowest flow sets how long a training job takes. IP Infusion develops OcNOS-DC, the network operating system that runs these switches, and ships it as complete, validated systems on 51.2 Tbps Broadcom Tomahawk 5 and 25.6 and 12.8 Tbps platforms.

51.2 TbpsTomahawk 5 platforms
800G64 ports per switch
RoCEv2lossless Ethernet transport
40+validated platforms
Definition

What makes a data center switch an AI switch

A switch earns the AI label when it carries GPU traffic at the speed, scale and loss rate a training cluster needs. Five properties set it apart. The network these switches form is covered in What is an AI fabric; this page is about the switch itself.

  • Port speed and capacity. GPU NICs attach at 400G and 800G, and the switch port has to match them. A 51.2 Tbps switch provides 64 ports at 800G; a 25.6 Tbps switch provides 64 ports at 400G.
  • Radix. The number of ports per switch sets how many GPUs one leaf can serve and how many leaves one spine can join, which decides how many tiers the fabric needs. Fewer tiers means fewer hops, fewer optics and less power.
  • Lossless transport. GPUs exchange data with RDMA over Converged Ethernet (RoCEv2), which slows sharply when packets drop. The switch needs Priority Flow Control (PFC) to pause a queue before it overflows and Explicit Congestion Notification (ECN) to warn senders early, working with DCQCN congestion control on the NIC.
  • Even load. Collective operations such as AllReduce produce a few very large flows. Static hashing can put two of them on one link while another link sits idle. Dynamic load balancing moves traffic toward less loaded paths.
  • Visibility. Streaming telemetry reports queue depth, PFC pauses and ECN marks as they happen, so operators can find the one link that is slowing a job.
Where they sit

Where AI switches sit in an AI network fabric

An AI data center runs four networks side by side, and each asks something different of its switches. The GPU back-end fabric is what most people mean by an AI switch.

GPU back-end fabric

Connects GPU NICs across servers for training and inference. Built from rail-optimized leaves and a Clos spine at 400G or 800G and run lossless end to end. This is where port speed and radix matter most.

Storage fabric

Feeds training data to the GPU servers and writes checkpoints back. Often RoCEv2 as well, sized for sustained bandwidth rather than GPU count.

Front-end network

Connects the cluster to users, orchestration and the rest of the data center. A conventional leaf-spine, often with EVPN-VXLAN to keep tenants apart in a GPU cloud.

Out-of-band management

A separate 1G network for consoles, server management controllers and switch management, so operators can reach every device when the data planes are busy or down.

Inside the back-end fabric, the same switch model can serve as a rail leaf, a spine or a super-spine. A rail-optimized design connects each GPU NIC in a server to a different leaf, so same-rank GPU traffic stays one hop away. The topology sets the role; AI fabric topologies and rail-optimized network cover those designs.

AI switches are bought by neocloud and GPU cloud providers, large enterprise data centers and MSPs, and service providers adding GPU capacity in their own data centers.

Sizing

800G or 400G: sizing an AI switch by radix

The port count of the switch decides how far a fabric grows before it needs another tier. In a non-blocking two-tier leaf-spine, each leaf splits its ports evenly between GPUs and spines, and each spine port connects one leaf. A switch with k ports then supports k squared divided by two endpoints in two tiers, and k cubed divided by four in three tiers.

Switch capacityPorts at full rateTwo tiers, non-blockingThree tiers, non-blocking
51.2 Tbps64 at 800G2,048 endpoints at 800G65,536 endpoints at 800G
25.6 Tbps64 at 400G2,048 endpoints at 400G65,536 endpoints at 400G
12.8 Tbps32 at 400G512 endpoints at 400G8,192 endpoints at 400G
  • These are arithmetic ceilings, not product scale figures. Real designs move them with oversubscription, rails per server, optics reach and rack power.
  • 800G at the same radix doubles the bandwidth per GPU without adding a tier. That is why back-end fabrics built around 800G NICs use 51.2 Tbps switches.
  • IP Infusion reference designs for rail-optimized and 3-stage Clos fabrics, sized in real port counts, are on AI fabric topologies. The choice between designs is walked through in how to choose an AI networking fabric.
OcNOS platforms

AI data center switches validated for OcNOS-DC

OcNOS-DC is the IP Infusion network operating system for data center and AI fabrics. It runs on open switches built on Broadcom silicon, and every platform below is on the OcNOS Hardware Compatibility List, which covers 40+ validated platforms in all.

CapacityBroadcom ASICValidated platformsTypical place in an AI data center
51.2 TbpsTomahawk 5 (BCM78900)Edgecore AIS800-64D, UfiSpace S9321-64E, UfiSpace S9321-64EO800G back-end leaf and spine; the S9321-64EO also carries coherent DCI between sites
25.6 TbpsTomahawk 4 (BCM56990)Edgecore AS9736-64D (DCS520)400G back-end fabric, storage or front-end
12.8 TbpsTomahawk 3 (BCM56980)Edgecore AS9716-32D (DCS510)400G spine for smaller fabrics and front-end networks
12.8 TbpsTrident 4 (BCM56880)Edgecore AS9726-32DB (DCS240), UfiSpace S9300-32D400G leaf for front-end and storage; IPoDWDM capable

What OcNOS-DC brings to the switch

  • Lossless RoCEv2: PFC, including PFC over Layer 3 routed underlays, PFC deadlock detection and recovery, ECN with DCQCN tuning, and DCBX and ETS for per-class bandwidth.
  • Load balancing: Dynamic Load Balancing (DLB) for flowlet-aware adaptive routing, and Global Load Balancing (GLB) for fabric-wide path optimization in OcNOS 7.1.
  • Multi-tenant GPU clouds: EVPN-VXLAN keeps tenants apart on a shared fabric. On Trident 3 X4, X5 and X7 and Trident 4 platforms, ECN and PFC over VXLAN also signal congestion across the overlay.
  • Operations: streaming telemetry over gNMI and OpenConfig, plus NETCONF, YANG and zero-touch provisioning.
  • Licensing and platform support: DLB, PFC deadlock detection and recovery, and ECN over VXLAN are OcNOS-DC PLUS tier features, and support varies by platform. The OcNOS Feature Matrix lists 700+ features by platform and tier.
  • Management network: 1G Trident 3 X2 platforms (Edgecore AS4625-54T, Celestica DS1000 and UfiSpace S6301-56ST, 120 Gbps each) run the out-of-band network on the same NOS.
Benefits and trade-offs

What open AI switches give you, and what they ask of you

Open switches with a separate network operating system change where the choices sit and where the work sits.

What you gain

  • Hardware choice. OcNOS-DC runs on switches from more than one hardware maker, so each plane can be chosen on capacity, optics and lead time.
  • One NOS across planes. Back-end, storage, front-end and management networks run one operating system, one CLI and one automation model.
  • Standard Ethernet. A RoCEv2 fabric uses the same optics, cabling and tooling as the rest of the data center, and the GPU NIC is chosen separately from the switch.
  • A path to new silicon. Moving from 25.6 to 51.2 Tbps changes the switch, not the operating model.

What it asks of you

  • Integration. Lossless Ethernet only works when switch buffers, PFC, ECN and the NIC's congestion control are tuned together. Someone has to do that tuning and own it.
  • Support model. With hardware and software from different companies, the operator needs one clear owner when a fault crosses the line between them.
  • Validation burden. Optics, cables, NIC firmware and NOS releases each change over time, and every combination needs testing before production.

OcNOS Systems address these three points by shipping the switch and OcNOS-DC as one complete, validated system, ready to deploy, with the supported platform and feature combinations published in the Hardware Compatibility List and the Feature Matrix.

Compare

Three ways to buy AI data center switches

The switch silicon is often the same across the options. What differs is who picks the software, who integrates the pieces, and who carries the testing.

QuestionIntegrated switch and NOS from one makerSelf-integrated open switchValidated open system
Hardware choiceThe maker's own modelsAny open switch the buyer selectsOpen switches on the OcNOS HCL
Network operating systemTied to the hardwareChosen and installed by the buyerOcNOS-DC, validated on that platform
Who integrates and testsThe makerThe buyerIP Infusion validates the platform and NOS combination
GPU NIC choiceVaries by makerAny NIC that supports RoCEv2Any NIC that supports RoCEv2
Where the effort sitsLow integration effort, narrower choiceHighest integration effort, widest choiceIntegration done up front, choice within the HCL
Moving to new siliconWhen the maker releases itWhen the buyer has validated itWhen the platform is added to the HCL

AI Networking Switches FAQ

What is an AI switch?
An AI switch is an Ethernet data center switch built for the GPU back-end fabric. It offers 400G or 800G ports at high radix, supports lossless RoCEv2 with PFC and ECN, and spreads large flows across every path, because the slowest flow sets how long a training job takes. The term describes what the switch is used for and what it supports, not a separate class of hardware.
What is the best switch for AI?
The right switch depends on the GPU count, the NIC speed and how many fabric tiers you accept. For 800G NICs, a 51.2 Tbps switch with 64 ports at 800G has an arithmetic ceiling of 2,048 endpoints in a two-tier non-blocking fabric. For 400G NICs, a 25.6 Tbps switch gives the same radix at 400G. Beyond port speed, check PFC and ECN support, dynamic load balancing, streaming telemetry, and whether the network operating system is validated on that hardware.
Do I need 800G switches for AI, or is 400G enough?
Match the switch to the NIC. Clusters with 400G NICs are matched by 25.6 or 12.8 Tbps switches. Clusters built around 800G NICs need 800G switch ports, or each GPU gets half its network bandwidth, which in practice means a 51.2 Tbps switch. At the same radix, 800G doubles the bandwidth per GPU without adding a fabric tier.
What is a GPU fabric switch?
It is another name for an AI switch in the GPU back-end network: the leaf, spine or super-spine switches that connect GPU NICs across servers. Inside a server or rack, GPUs also talk over a separate scale-up interconnect. The GPU fabric switch carries the scale-out traffic between servers.
Do AI switches need deep buffers?
Usually not in the back-end fabric. Most AI back-end fabrics use high-radix switches with on-chip shared buffers and rely on PFC, ECN and load balancing to prevent loss. Deep-buffer switches are more common at the edge of the network, such as data center interconnect and service provider routers, where traffic arrives over long or mismatched links.
Can an AI network fabric use Ethernet instead of InfiniBand?
Yes. RoCEv2 carries RDMA over standard Ethernet, and with PFC, ECN and load balancing an Ethernet fabric runs lossless for GPU traffic. Ethernet also lets one operations team, one set of optics and one toolset cover the back-end and front-end networks.
Which switches does OcNOS support for AI fabrics?
For the GPU back-end fabric, OcNOS-DC runs on 51.2 Tbps Broadcom Tomahawk 5 switches (Edgecore AIS800-64D, UfiSpace S9321-64E and UfiSpace S9321-64EO) and a 25.6 Tbps Tomahawk 4 switch (Edgecore AS9736-64D). OcNOS-DC is also validated on 12.8 Tbps Tomahawk 3 and Trident 4 switches (Edgecore AS9716-32D, Edgecore AS9726-32DB and UfiSpace S9300-32D). The full list of 40+ validated platforms is in the OcNOS Hardware Compatibility List.
Related

Keep reading on AI fabrics and open networking

What is an AI fabric?

The GPU back-end network these switches build.

AI fabric architecture

Every plane of an AI data center, and its hardware.

OcNOS AI Fabric

The complete 800G AI fabric on OcNOS-DC.

InfiniBand vs Ethernet

The two transports compared for AI.

What is a white box switch?

The open hardware model behind AI switches.

What is a network operating system?

The software that turns silicon into a fabric.

What is network disaggregation?

Buying the switch and its software separately.

Open networking

Open hardware, open software, and where IP Infusion fits.

What is lossless Ethernet?

PFC, ECN and DCQCN, and how they prevent loss.

Choosing switches for a GPU cluster? Start from the GPU count.

Tell us the GPU count, the NIC speed and the oversubscription you can accept. An IP Infusion engineer will size the leaf, spine and super-spine stages on validated switches running OcNOS-DC.