How to choose an AI networking fabric
Your GPUs run at their rated throughput only if the fabric under them keeps up. This is a vendor-neutral framework for scoring an AI fabric on the six criteria that decide it, before you sign.
Read the fabric as one decision, made of six
An AI networking fabric is bought once and lived with for years, so score it on how six properties fit the way you will run the cluster, not on a single headline number. Three kinds of fabric sit on the table:
- Closed single-vendor stacks bundle the NIC, switch, and scheduler into one SKU and set the trade-offs for you.
- Community and open-source NOS options give you code but leave integration and support on your desk.
- Open commercial NOS options keep hardware choice while carrying a support contract.
The criteria below score any of them on the same terms.
Performance and losslessness
AI collectives use RDMA, and RDMA needs a lossless substrate. RoCEv2 delivers that on Ethernet, but only with the full congestion toolkit, not a checkbox. A fabric that pauses badly or drops under load is the most common cause of stalled training throughput, so settle this criterion first.
What to ask
- Do you deliver the lossless Ethernet fabric RoCEv2 requires, including PFC over L3 with DCBX and LLDP?
- Is ECN marking present, and is there a dynamic variant that adapts thresholds under load?
- Is there PFC deadlock detection and recovery, plus WRED and buffer tuning?
- How is tail latency held down when a link gets hot during a collective?
How to evaluate
Ask for the exact congestion features and confirm they run on the silicon you will buy. A credible answer names PFC (including over L3, with DCBX and LLDP), ECN and Dynamic ECN, ETS, WRED, PFC deadlock detection and recovery, and adaptive load balancing that rebalances flows off hot links. A static hash can leave large flows colliding on a subset of links, while adaptive routing holds effective fabric utilisation in the 90% range on Tomahawk 4 and Tomahawk 5, so treat it as a throughput feature, not a nicety. See the RoCE 与 InfiniBand 对比 comparison for how the lossless model differs from InfiniBand credit flow.
Scale and topology fit
The right topology is the smallest non-blocking design that keeps every GPU's link saturated during a collective, with a clean path to the next size up. A topology that cannot grow forces a forklift later, so size for where the cluster is going, not where it starts.
What to ask
- Is the GPU plane 1:1 non-blocking, and where does oversubscription begin?
- What is the largest cluster this design reaches before a tier has to be added?
- Does it support rail-optimized wiring and a 3-stage Clos for multi-pod scale?
How to evaluate
Map candidates against concrete GPU counts. On radix-64 Tomahawk 5 switches, a 2-tier leaf-spine reaches about 2,048 GPUs at 1:1 non-blocking with one 800G fabric port per GPU, and scales toward 8,000+ GPUs when 800G ports break out to multiple NICs. A 3-stage Clos with a super-spine tier extends to 16,000+ GPUs, up to about 65,536 at the fat-tree limit. Use rail-optimized wiring for single-pod locality and a Clos for multi-pod scale; the 参考拓扑 show the port maths, and the AI 网络设计套件 gives a first cut.
Openness and lock-in
A closed AI fabric ties the NIC, switch, and scheduler to one supplier, which sets your pricing and roadmap for you. An open approach runs a commercial NOS on standard Ethernet silicon from more than one hardware vendor, so you keep hardware choice and can second-source. This is where multi-year cost and negotiating leverage are decided.
What to ask
- Can the switch, the NOS, and the optics each be sourced independently?
- Is the fabric built on published, standards-based Ethernet or a proprietary transport?
- Is there more than one qualified hardware platform for the same software?
How to evaluate
Look for a validated hardware list with more than one vendor behind it. Open commercial NOS options run the same software across Broadcom Tomahawk 5 platforms from different makers, for example Edgecore AIS800-64D and UfiSpace S9321-64E, keeping a second source real rather than theoretical. Contrast that with closed single-vendor NIC, switch, and scheduler SKUs where the parts only work together. The useful question is not one product versus another but how many suppliers you can hold accountable for the same fabric.
Operations, automation, and telemetry
A fabric that is hard to bring up or blind in production costs more than its price tag. The operational model, how you provision, observe, and tune, is a first-class buying criterion, not an afterthought once the hardware lands.
What to ask
- Is there streaming telemetry over gNMI with OpenConfig models?
- Is zero-touch provisioning (ZTP) supported for fabric bring-up at scale?
- Does it automate with standard tools such as Ansible, and signal QoS with DCBX?
How to evaluate
Ask to see the telemetry and automation in the candidate's own docs, not a slide. A strong answer covers gNMI and OpenConfig streaming telemetry into your observability stack, ZTP for provisioning, DCBX and LLDP for lossless signalling, and Ansible-based configuration. The point is closed-loop tuning during bring-up: per-queue counters and buffer state you can watch while you shake out PFC and ECN thresholds, not a black box you set once and hope.
Support and accountability
When a training run stalls at 3 a.m., the question is who owns the fix. A multi-vendor build with no single accountable contract turns into finger-pointing between the NIC, switch, optics, and software vendors. The remedy is one contract, not one closed stack.
What to ask
- Is there a single support contract covering software and the validated hardware?
- Is there one TAC and one SLA, or several to coordinate?
- Who validated the specific hardware and software combination you are buying?
How to evaluate
Separate accountability from lock-in: you want single-vendor accountability without single-vendor hardware. The strong pattern is open hardware choice paired with one support agreement, one contract covering the software plus the validated hardware, one TAC, one SLA, so a fault is owned end to end even though you sourced the parts openly.
Standards and roadmap
AI fabric 传输发展迅速。Ultra Ethernet Consortium(UEC)发布了面向 AI 与 HPC 的开放多路径以太网传输,其端点行为运行在支持 UEC 的 NIC 上,因此请区分各厂商当前在生产环境中运行的能力与其规划中的能力。
What to ask
- 该厂商是否为 Ultra Ethernet Consortium 成员?
- 该 fabric 当前在生产环境中运行什么,基于哪款芯片?
- What is on the near-term roadmap for load balancing and congestion control?
How to evaluate
区分已交付的能力与规划中的能力。IP Infusion 是 Ultra Ethernet Consortium 的成员,OcNOS-DC 如今已在 Broadcom Tomahawk 4 和 Tomahawk 5 交换机上,于生产环境中运行基于 PFC、ECN、DCQCN 和亚毫秒级 DLB 的无损 RoCEv2。关于路线图,请索要具体项目:OcNOS 7.1 新增 Global Load Balancing(GLB)与基于时延的 ECN(7.1.0),链路级重试、基于信用的流控和数据包裁剪已列入路线图。请参阅 InfiniBand 与以太网对比 for how the open path compares.
The buyer checklist
Run every candidate fabric, closed single-vendor stack, community NOS, or open commercial NOS, through the same list. If a box cannot be ticked, ask why before you decide it does not matter.
- Lossless substrate. RoCEv2 with PFC (including over L3, DCBX, LLDP), ECN and Dynamic ECN, ETS, WRED, and PFC deadlock detection and recovery.
- Adaptive load balancing. Dynamic load balancing that rebalances flows off hot links, not just static ECMP hashing.
- 1:1 non-blocking GPU plane. No oversubscription on the GPU fabric, sized for your target GPU count with headroom to the next tier.
- Topology fit. Rail-optimized single-pod and 3-stage Clos multi-pod designs on the silicon you are buying.
- 硬件选择。 More than one validated hardware vendor for the same NOS, so a second source is real.
- Open transport. Standards-based Ethernet, not a proprietary NIC, switch, and scheduler bundle.
- Telemetry and automation. gNMI and OpenConfig streaming telemetry, ZTP, DCBX, and Ansible.
- One accountable contract. Single support agreement across software and validated hardware, one TAC, one SLA.
- Standards roadmap. 厂商是 Ultra Ethernet Consortium 成员,并明确列出其负载均衡与拥塞控制的路线图项目。
Where open Ethernet with OcNOS-DC fits
Read the criteria above and pick whatever scores best for your cluster. For buyers who want open hardware choice with a single accountable contract, here is how IP Infusion lines up against the same checklist, stated plainly so you can compare it with the alternatives.
Production RoCEv2 fabric
OcNOS-DC delivers the lossless Ethernet fabric RoCEv2 requires on Broadcom Tomahawk 5: PFC over L3 with DCBX and LLDP, ECN and Dynamic ECN, ETS, WRED, PFC deadlock detection and recovery, and DLB with reactive path rebalance.
2,048 to 16,000+ GPUs
A 2-tier TH5 pod reaches about 2,048 GPUs at 1:1 non-blocking and toward 8,000+ with breakout; a 3-stage Clos extends to 16,000+ and up to about 65,536 at the fat-tree limit.
Multi-vendor hardware
One NOS across Tomahawk 5 platforms from more than one maker, including Edgecore AIS800-64D and UfiSpace S9321-64E and S9321-64EO with 400G ZR+, so a second source stays real.
Open telemetry and automation
gNMI and OpenConfig streaming telemetry, ZTP, DCBX, and Ansible, with RTAG7 hashing and buffer tuning for closed-loop bring-up.
OcNOS Systems:一份合同、一个 TAC、一份 SLA
借助 OcNOS Systems,一份 IP Infusion 协议即可涵盖 OcNOS 软件与经过验证的硬件,让您获得单一厂商的责任归属,而无单一厂商锁定。
UEC 成员,路线图明确
IP Infusion 是 Ultra Ethernet Consortium 的成员。GLB 与基于时延的 ECN 已列入面向 OcNOS-DC 的 OcNOS 7.1 版本系列。
AI fabric buyer guide FAQ
What should I evaluate first when choosing an AI fabric?
Is Ethernet or InfiniBand better for AI?
How many GPUs do I need to plan for?
How do I avoid vendor lock-in?
What does UEC mean for my purchase?
Do I need one vendor for hardware and software?
Scoring fabrics for your cluster? We'll run the checklist with you.
预约架构评审 →用 OcNOS 设计整张 AI fabric
从商业论证到端口数量测算,无论您进行到哪一步,都可从此接续。
深入了解,随身带走。
两份简短的技术下载资料,比本页更深入:完整的 OcNOS-DC 数据表和无损 800G AI fabric 架构。
OcNOS-DC 数据手册
表单简短。提交后您的 PDF 将立即在新标签页中打开。
✓ 正在新标签页中打开您的 PDF……
如果未能自动打开,请使用下方链接。
OcNOS 800G 无损 AI Fabric
表单简短。提交后您的 PDF 将立即在新标签页中打开。
✓ 正在新标签页中打开您的 PDF……
如果未能自动打开,请使用下方链接。