Balanced Bandwidth per GPU: Sizing Cluster Networks for LLM Training

NoraLin 12 2026-10-07 03:23:00 Edit

Buy GPUs by the count the model needs; buy the network by the arithmetic the communication demands. Most cluster disappointments are mismatched pairs — fast accelerators behind a fabric that turns every synchronization into a queue. The math that prevents this is not deep: three inputs, one derivation, and one acceptance test that measures what the vendor quoted. This page is that arithmetic.

Prerequisites: Model, Parallelism, Traffic

Three inputs derive the requirement before any fabric is chosen: the model's communication volume (parameter count and gradient size set the bytes each synchronization moves), the parallelism strategy (data-parallel all-reduce dominates some tiers, tensor-parallel traffic stays within NVLink domains, pipeline-parallel sends point-to-point), and the target step time — because per-GPU bandwidth is an output of this arithmetic, and engineering guidance already frames the envelope: multi-node training wants roughly 200 Gbps to 1 Tbps per node to scale efficiently.

InputWhat it determinesWhere it comes from
Model communication volumeBytes per synchronization eventParameter count and gradient size
Parallelism strategyWhich traffic tier crosses the fabricYour training architecture
Target step timeHow much of each step communication may consumeYour throughput objective
The envelope checkWhether the derivation is plausible~200 Gbps–1 Tbps per node guidance

The parallelism mapping deserves the most care because it decides where traffic lives: data-parallel training moves gradients across the fabric at every synchronization — the all-reduce pattern whose bandwidth consumption is well modeled — while tensor-parallel traffic mostly stays inside the NVLink domain of a node and pipeline-parallel moves point-to-point at stage boundaries. Same model, three different fabrics depending on the strategy.

Size the Fabric: NICs, Bisection, Oversubscription

Sizing runs three steps: pick the NIC per GPU from the traffic tier (the communication-heavy tiers get full-rate adapters, in the documented pattern of one 400G-class port per accelerator), decide bisection bandwidth for the tiers that synchronize (near-full bisection where all-reduce crosses the fabric — bisection is the cross-section capacity between halves, the number that decides whether synchronization waits), and place deliberate oversubscription only where traffic tolerates it — because a fabric oversubscribed where all-reduce lives converts every step into a wait, and research quantifies exactly that scaling penalty.

  1. NIC per GPU: full-rate adapters (400G-class per accelerator as the common pattern) for tiers whose synchronization crosses the fabric; lighter tiers may share.
  2. Bisection for sync tiers: near-full bisection bandwidth where all-reduce dominates — the design parameter specialized training fabrics are built around.
  3. Deliberate oversubscription: convergence above 1:1 only on tiers with tolerant traffic, decided per tier and written down, never applied fabric-wide by accident.
  4. Placement: communication-heavy ranks bound to the nearest physical topology, the pattern rail-aligned and topology-aware fabrics implement.

This is where the balanced-design argument lands: environments engineered with matched compute, network, and storage scale — such as OneSource Cloud's dedicated clusters on a Spine-Leaf RoCE v2 fabric with hardware offload and zero-loss priority flow control — exist precisely because the derivation above, done honestly, produces a fabric specification that generic commodity deployments do not meet.

Verify: Measure Delivered Bandwidth Before Acceptance

Verification measures delivered bandwidth, not datasheet bandwidth: run all-reduce bus-bandwidth tests at the sizes your training actually issues, at the node counts you will actually run, before accepting the cluster — the all-reduce consumption analysis supplies the physics to interpret results, and a fabric that measures below its design at acceptance will measure below it during every training run it ever hosts.

  • Test at real message sizes: all-reduce at the byte counts your gradients actually produce, not a synthetic maximum.
  • Test at real node counts: bisection behavior changes with scale — test the fleet you will run.
  • Threshold from design: the acceptance bar is the derivation from step one, written before the test.
  • Re-test on topology change: any fabric change re-runs the suite; yesterday's number does not cover today's topology.

The acceptance test is also your negotiating position: a cluster that delivers 60% of designed bus bandwidth at your sizes is a cluster whose every training step will pay that gap, and the difference between quoted and delivered is easiest to fix before the final payment, on the vendor's initiative rather than yours.

FAQ

How do I calculate network bandwidth per GPU for training?

Start from the traffic, not the catalog: gradient bytes per step from parameter count, multiplied by synchronization frequency, divided across the tiers your parallelism strategy assigns to the fabric — then check the product against the documented envelope of roughly 200 Gbps to 1 Tbps per node; when your derived number sits below the envelope, oversubscription is a choice, and when it sits above it, the fabric is the bottleneck before the first job runs.

When is an oversubscribed fabric acceptable for training?

When the oversubscribed tier carries tolerant traffic: tiers where communication stays inside NVLink domains or moves point-to-point at pipeline pace tolerate convergence above 2:1, while tiers hosting all-reduce want near-full bisection — the mistake is applying one ratio to the whole fabric instead of per tier, because oversubscription placed where synchronization lives converts every step into queueing.

What test proves the fabric delivers its design bandwidth?

All-reduce bus bandwidth at your real message sizes and node counts, run before acceptance and re-run on topology changes: datasheet line rates describe the silicon, busbw at your sizes describes the path your gradients will take, and the gap between the two is the wait your training will do every step — accept the measured number, not the quoted one.

Previous: Flat Rate Billing for AI GPU Cloud
Next: How Much Spare GPU Capacity Inference Needs for Failover
Related Articles