It is tempting to think of a data center network primarily in terms of connectivity to the outside world — the bandwidth available to internet exchanges, cloud regions, and end users. For AI training clusters, that "north-south" traffic is often secondary to a far more demanding requirement: the "east-west" fabric connecting thousands of GPUs to one another within the cluster itself.
Why GPU-to-GPU Bandwidth Dominates the Design Problem
Training a large AI model is fundamentally a distributed computing problem, with thousands of accelerators exchanging enormous volumes of data — gradients, model parameters, intermediate activations — at extremely high frequency throughout the training process. The speed of this internal communication directly determines how efficiently the cluster's raw compute capacity translates into useful training throughput. A cluster with world-class GPUs but an inadequate network fabric can sit substantially underutilised, with accelerators idling while waiting for data to arrive from other nodes.
This has made network fabric design a first-order architectural decision for AI infrastructure, rather than a downstream implementation detail handled after compute and cooling decisions are finalised.
What a Purpose-Built AI Fabric Actually Requires
- Extremely high bandwidth per connection, often far exceeding what general-purpose enterprise or cloud networking requires
- Very low, predictable latency, since synchronised training jobs are often only as fast as their slowest communication path
- Network topologies — such as fat-tree or other non-blocking architectures — specifically designed to minimise congestion across many-to-many communication patterns typical of distributed training
- Careful physical layout planning, since cable lengths and switch placement directly affect achievable latency at the scale of thousands of connected accelerators
A data center's network fabric is no longer just plumbing for data — for AI training clusters, it is as much a determinant of effective compute capacity as the accelerators themselves.
Facility Design Implications
Because network fabric performance is sensitive to physical layout, AI factory floor plans increasingly need to be designed with network topology as a primary input, not an afterthought fitted around an electrical and mechanical layout decided independently. This can mean trade-offs between optimal network topology and optimal cooling or power distribution layout — trade-offs that require multidisciplinary coordination from the earliest design stages rather than being resolved late in detailed engineering.
It also means that retrofitting an existing, non-AI-optimised facility to support frontier-scale training clusters can be considerably more constrained than building a purpose-designed AI factory, since the physical layout decisions that most affect network performance are often baked into a building's structure from the outset.
Networking Equipment Has Its Own Supply Chain Dynamics
Like electrical equipment, high-performance networking hardware for AI clusters is subject to its own demand pressures, with leading switch and networking platform vendors managing substantial order backlogs as cluster build-outs accelerate across the industry. Developers and operators planning large AI clusters increasingly need to factor networking equipment lead times into overall project scheduling, alongside the more widely discussed power and cooling equipment constraints — a cluster with secured power and cooling capacity but an unfilled networking equipment order is no more operational than one missing any other critical system.
Connectivity Beyond the Cluster Still Matters
While east-west fabric design dominates training cluster architecture, north-south connectivity to the broader network remains important — particularly for inference workloads serving external users, and for facilities that need to move large training datasets or model checkpoints to and from other locations. A well-designed AI facility addresses both requirements deliberately, rather than over-optimising for one at the expense of the other.
DATAPERT's technical advisory work integrates network fabric design into the broader data center development process from the earliest planning stages. Explore our technology integration capabilities or start a project to discuss network architecture for an AI-ready facility.
