Neocloud · AI training and inference

Networking for the neocloud: one fabric for AI training and inference

A neocloud runs training and inference on the same infrastructure. IP Infusion delivers the open network for both: OcNOS on validated open hardware, lossless RoCEv2 for GPU clusters, distributed inference at the edge, and open transport between sites, from one control plane and one support contract.

1 fabrictraining and inference
up to 800GEthernet AI fabric
40 to 60%lower hardware cost
600+operator networks
Two workloads, one business

Training and inference pull the network in opposite directions

Every neocloud serves both. Training fills a few large clusters with steady, synchronized traffic. Inference spreads across many sites and spikes without warning. The fabric has to do justice to both, or the workloads that pay the bills go somewhere else.

Training vs inference, across the network metrics that shape the fabric.
What the network sees Trainingthe AI factory Inferencethe real-time service
FootprintA few large GPU clustersMany regional and edge sites
TrafficBulk, synchronized GPU-to-GPU collectivesSmall, latency-critical requests
Load profileSustained, near full GPU utilization for long runsBursty and elastic, spikes in seconds
The network has to beLossless under sustained bandwidthLow-latency and redundant at every layer
Failure toleranceTolerant, jobs checkpoint and resumeLow, tight SLAs on every request
The network decision that sets your margin

One fabric, or two?

Every neocloud answers the same question: how many networks does it build to serve training and inference? The answer shapes the cost base for the life of the buildout.

Option A · two networks

A training network, plus a second one for inference

One fabric for training, a different one bolted on for inference. Two designs, two spares pools, two operations teams, two upgrade cycles. Capital and operating cost are duplicated before the first customer arrives, and the margin has to absorb all of it.

Duplicated capex and opex, margin under pressure
Option B · one OcNOS fabric

Training and inference on one control plane

Both workloads run on the same OcNOS network. One image, one control plane, one support contract. The training fabric and the distributed inference sites are the same platform, sized and licensed per role, so the operator captures training and every inference workload that follows on the infrastructure it already owns.

One fabric, cost that scales with the business
One OcNOS fabric for a neocloud: a leaf-spine GPU training cluster running lossless RoCEv2 at up to 800G on the left, distributed inference sites on the right, connected by open IP over DWDM transport, all under one OcNOS control plane.
One OcNOS fabric for training and inference: a lossless RoCEv2 GPU training cluster, distributed inference sites, and open transport between them, under one control plane.
The training fabric

A lossless Ethernet fabric for the GPU clusters

Training moves bulk, synchronized traffic between GPUs, so the fabric has to stay lossless under load. OcNOS runs the RoCEv2 toolkit on open merchant silicon at up to 800G, so the AI factory gets a high-radix Ethernet fabric without a single-vendor stack.

Lossless RoCEv2

Priority flow control, ECN, and DCQCN keep GPU collective traffic out of packet loss, with adaptive load balancing to spread it across the fabric.

Up to 800G on open silicon

High-radix Ethernet on validated open hardware from multiple vendors carries the training fabric, with the headroom to scale the cluster out.

Ready for the GPU stack

The fabric carries the GPU collective libraries (NCCL, RCCL, oneCCL) that training jobs run on, so the network is not the bottleneck in the AllReduce.

One image, every tier of scale
Rack
GPU servers + leaf
The unit you add capacity in: GPU servers behind a top-of-rack leaf.
Pod
Leaf-spine block
A non-blocking leaf-spine block, the repeatable building block of the fabric.
Cluster
Up to 4,096 GPUs
A full training cluster on one lossless RoCEv2 fabric.
Multi-cluster
16,384-GPU designs
Multiple clusters plus distributed inference sites, under one control plane.

The same OcNOS image runs every tier, so the fabric scales out with the cluster instead of being redesigned at each step.

Distributed inference

Inference lives at the edge, and needs the network to follow

Inference spreads across many regional and edge sites, scales up and down quickly, and holds tight latency targets. This is where a neocloud needs more than a data center fabric: it needs transport between sites and assurance across all of them. IP Infusion covers the whole path.

Data center fabric at each site

An EVPN-VXLAN leaf-spine fabric on open switches serves each inference site, with the elasticity to add and remove capacity as demand moves.

Open transport between sites

Low-latency transport ties the sites together. IP over DWDM with coherent ZR and ZR+ collapses the optical layer onto the router for site-to-site capacity.

Assurance across every site

IP Maestro gives one view of the whole footprint, so an operator holds inference service levels across many sites and N+1 designs.

Open hardware, lower cost

A cost base a neocloud can defend

A neocloud competes on price and time to scale, so the network cannot be a proprietary tax. Open networking puts the operator in control of hardware, software, and lead times, with one vendor still owning the software and support.

Multi-vendor supply chain

Switches from Edgecore, UfiSpace, and others run the same OcNOS image, so the operator is never tied to one vendor for hardware or lead times.

Lower hardware cost

OcNOS on open merchant-silicon switches carries the AI fabric at 40 to 60% lower hardware cost than proprietary platforms at comparable port speeds and radix.

One vendor owns the fix

Hardware and software refresh on independent cycles, but one support contract covers the complete system, so one team owns the fix.

Proven in production

The open network is already running at scale

The neocloud fabric is not a lab exercise. It runs the same OcNOS that carries production traffic for operators around the world, on the same open hardware.

600+
operator networks run on OcNOS across service providers, data centers, and internet exchanges.
60+
countries where OcNOS carries live production traffic today.
40+
validated open hardware platforms run OcNOS from one image.
FAQ

Neocloud networking, answered

What is a neocloud?
A neocloud is an AI-first cloud provider built around dense GPU compute for training and inference, rather than the general-purpose services of a traditional hyperscaler. Neoclouds compete on raw accelerator performance, fast time to scale, and cost, so the network that carries GPU traffic is a core part of the economics, not an afterthought.
Do neoclouds use InfiniBand or Ethernet?
Both are deployed, and Ethernet is where the industry is heading as 800G ramps. Ethernet with RoCEv2 carries GPU collective traffic on open merchant silicon, so a neocloud gets a lossless AI fabric without a single-vendor stack. OcNOS runs the full RoCEv2 toolkit, priority flow control, ECN, and adaptive load balancing, on validated open hardware.
Should a neocloud build one network or two for training and inference?
Training and inference are opposite workloads: training is centralized and bandwidth-heavy, inference is distributed, bursty, and latency-critical. Building two separate networks duplicates capital and operating cost. Running both on one OcNOS fabric with one control plane lets a neocloud serve training and every inference workload from the same infrastructure, which is where the operating margin improves.
When does a neocloud still need two separate networks?
Some operators separate training and inference on purpose, and there are good reasons to: strict isolation between tenants, distinct operational domains for different teams, a hard performance boundary, or an existing InfiniBand island that already carries training. OcNOS runs either way. The point is that separation should be a deliberate design choice, not a cost the architecture forces on you. When one team can run both workloads on one control plane, one fabric sized and licensed per role is the default that protects margin.
What does distributed inference need from the network?
Inference spreads across many regional and edge sites, scales up and down quickly, and holds tight latency targets, so the network needs elastic capacity, redundancy at every layer, and low-latency transport between sites. IP Infusion covers the full path: the data center fabric, the transport and IP over DWDM between sites, and IP Maestro for assurance.
How does open networking lower neocloud cost?
A neocloud buys switches and software on independent cycles from a multi-vendor supply chain, so it is not tied to one vendor for hardware, software, or lead times. OcNOS on open merchant-silicon switches carries the AI fabric at a lower hardware cost than proprietary platforms, and one vendor still owns the software and support under a single contract.
Datasheet

Take the OcNOS-DC datasheet with you

A short, technical download that goes further than this page: the full OcNOS-DC datasheet.

Design your neocloud fabric with one control plane

Tell us the GPU scale and the sites you plan to serve, and an IP Infusion engineer will help you design one open fabric for training and inference.