Networking for the neocloud: one fabric for AI training and inference
A neocloud runs training and inference on the same infrastructure. IP Infusion delivers the open network for both: OcNOS on validated open hardware, lossless RoCEv2 for GPU clusters, distributed inference at the edge, and open transport between sites, from one control plane and one support contract.
Training and inference pull the network in opposite directions
Every neocloud serves both. Training fills a few large clusters with steady, synchronized traffic. Inference spreads across many sites and spikes without warning. The fabric has to do justice to both, or the workloads that pay the bills go somewhere else.
| What the network sees | Trainingthe AI factory | Inferencethe real-time service |
|---|---|---|
| Footprint | A few large GPU clusters | Many regional and edge sites |
| Traffic | Bulk, synchronized GPU-to-GPU collectives | Small, latency-critical requests |
| Load profile | Sustained, near full GPU utilization for long runs | Bursty and elastic, spikes in seconds |
| The network has to be | Lossless under sustained bandwidth | Low-latency and redundant at every layer |
| Failure tolerance | Tolerant, jobs checkpoint and resume | Low, tight SLAs on every request |
One fabric, or two?
Every neocloud answers the same question: how many networks does it build to serve training and inference? The answer shapes the cost base for the life of the buildout.
A training network, plus a second one for inference
One fabric for training, a different one bolted on for inference. Two designs, two spares pools, two operations teams, two upgrade cycles. Capital and operating cost are duplicated before the first customer arrives, and the margin has to absorb all of it.
Training and inference on one control plane
Both workloads run on the same OcNOS network. One image, one control plane, one support contract. The training fabric and the distributed inference sites are the same platform, sized and licensed per role, so the operator captures training and every inference workload that follows on the infrastructure it already owns.

A lossless Ethernet fabric for the GPU clusters
Training moves bulk, synchronized traffic between GPUs, so the fabric has to stay lossless under load. OcNOS runs the RoCEv2 toolkit on open merchant silicon at up to 800G, so the AI factory gets a high-radix Ethernet fabric without a single-vendor stack.
Lossless RoCEv2
Priority flow control, ECN, and DCQCN keep GPU collective traffic out of packet loss, with adaptive load balancing to spread it across the fabric.
Up to 800G on open silicon
High-radix Ethernet on validated open hardware from multiple vendors carries the training fabric, with the headroom to scale the cluster out.
Ready for the GPU stack
The fabric carries the GPU collective libraries (NCCL, RCCL, oneCCL) that training jobs run on, so the network is not the bottleneck in the AllReduce.
The same OcNOS image runs every tier, so the fabric scales out with the cluster instead of being redesigned at each step.
Inference lives at the edge, and needs the network to follow
Inference spreads across many regional and edge sites, scales up and down quickly, and holds tight latency targets. This is where a neocloud needs more than a data center fabric: it needs transport between sites and assurance across all of them. IP Infusion covers the whole path.
Data center fabric at each site
An EVPN-VXLAN leaf-spine fabric on open switches serves each inference site, with the elasticity to add and remove capacity as demand moves.
Open transport between sites
Low-latency transport ties the sites together. IP over DWDM with coherent ZR and ZR+ collapses the optical layer onto the router for site-to-site capacity.
Assurance across every site
IP Maestro gives one view of the whole footprint, so an operator holds inference service levels across many sites and N+1 designs.
A cost base a neocloud can defend
A neocloud competes on price and time to scale, so the network cannot be a proprietary tax. Open networking puts the operator in control of hardware, software, and lead times, with one vendor still owning the software and support.
Multi-vendor supply chain
Switches from Edgecore, UfiSpace, and others run the same OcNOS image, so the operator is never tied to one vendor for hardware or lead times.
Lower hardware cost
OcNOS on open merchant-silicon switches carries the AI fabric at 40 to 60% lower hardware cost than proprietary platforms at comparable port speeds and radix.
One vendor owns the fix
Hardware and software refresh on independent cycles, but one support contract covers the complete system, so one team owns the fix.
The open network is already running at scale
The neocloud fabric is not a lab exercise. It runs the same OcNOS that carries production traffic for operators around the world, on the same open hardware.
Neocloud networking, answered
What is a neocloud?
Do neoclouds use InfiniBand or Ethernet?
Should a neocloud build one network or two for training and inference?
When does a neocloud still need two separate networks?
What does distributed inference need from the network?
How does open networking lower neocloud cost?
Take the OcNOS-DC datasheet with you
A short, technical download that goes further than this page: the full OcNOS-DC datasheet.
OcNOS-DC Datasheet
Quick form. Your PDF opens in a new tab immediately after submit.
✓ Opening your PDF in a new tab…
If it didn't open, use the link below.
Design your neocloud fabric with one control plane
Tell us the GPU scale and the sites you plan to serve, and an IP Infusion engineer will help you design one open fabric for training and inference.
Would you like us to reach out?
Leave your details and an IP Infusion engineer will help you design one open fabric for AI training and inference. We'll only be in touch if you'd like us to.