Two identical GPU clusters, the same model, the same code — and one delivers twice the training throughput of the other. The difference hides in a single ratio: Model Flops Utilization, the share of your fleet's theoretical compute that training actually converts into progress. When that ratio collapses, the instinct is to blame the hardware or buy more of it. The right move is arithmetic first: measure the ratio correctly, attribute the loss to its cause family, fix that family, and prove the recovery.
Prerequisites: Define and Measure Before Judging

Measure before judging: Model Flops Utilization is the ratio of achieved FLOPs to theoretical peak during training, and the number only means something once the convention is fixed — model-level FLOPs with or without attention, peak counted at what clock — because teams routinely compare MFU numbers computed on different conventions and conclude their cluster is broken when their arithmetic was; only a consistently-defined, per-rank-instrumented baseline can be diagnosed.
| Convention choice | Why it changes the number | The discipline |
| Model FLOPs with or without attention | Attention FLOPs add materially at long context | Pick one, write it down, never compare across |
| Peak at boost clock or base clock | Thermal reality keeps boost from being sustained | State the clock basis beside every number |
| Per-rank instrumentation | Fleet averages hide the slowest ranks that gate steps | Instrument per rank, not per job |
The metric itself is well-defined in engineering practice — achieved FLOPs over theoretical peak as the core measure of GPU usage efficiency during training. The diagnostics discipline is what most teams lack: the convention is the measuring instrument, and an instrument that changes between measurements cannot detect anything.
Attribute the Collapse to Its Cause Family
Attribution runs the timing split: separate each step's time into compute, communication, and stall per rank, and the collapse resolves into its family — inter-GPU imbalance (ranks finishing at different times while others wait, the cluster-level efficiency loss research isolates), single-GPU efficiency loss (kernels, memory, thermals), or stragglers and stalls (one slow node or link dragging every iteration) — and production evidence says this matters: identical clusters diverge to ~23% MFU on exactly these differences, so the fix follows the family, never the fashion.
- Split the step: per rank, how much time is compute, how much is collective communication, how much is stall — the three buckets every family lives in.
- Read the spread: ranks finishing unevenly with wait time dominating points at imbalance; uniformly slow ranks point at single-GPU efficiency; loss concentrated in bursts tied to one node or link points at a straggler.
- Check the cheap causes: thermals and health first — a GPU outside its thermal band loses efficiency and stretches training, and a failing node announces itself in the timing split before it dies.
Research on scalable parallelism treats cluster-level MFU as the product of explicitly separated factors — inter-GPU workload imbalance and single-GPU compute efficiency — and your diagnosis should too: a number that blends the two cannot be fixed, because the fixes belong to different owners.
Fix the Family and Verify the Recovery
Each family has its fix — imbalance yields to rebalanced partitioning and pipeline/memory-balance configurations, efficiency loss to kernel and thermal corrections, stragglers to draining the offender and topology-aware placement that keeps communication-heavy ranks close — and verification is the same in every case: the same MFU convention, measured before and after, with the delta attributed to the specific change; production systems hold ~55% MFU at twelve-thousand-GPU scale, which is the honest reference for what recovered looks like at enterprise scale.
- Imbalance fixes: repartition work across ranks, rebalance pipeline and memory configurations — the shaping work that makes ranks finish together.
- Efficiency fixes: kernel configurations, memory formats, thermal management — getting each GPU back to its own potential.
- Straggler fixes: drain the offending node, then placement that binds communication-heavy ranks to the nearest physical topology — the pattern topology-aware schedulers such as the OnePlus AI Orchestration Platform implement at the cluster layer, minimizing cross-switch and cross-NUMA delay for communication-dense ranks.
- Verification rule: same convention, before and after, delta attributed to the one change you made.
The before/after rule is what keeps recovery honest: production reference points — 55.2% at 12,288 GPUs in a full-stack engineering report — calibrate ambition, but only your own fixed-convention delta proves the fix, and a balanced dedicated environment from a provider such as OneSource Cloud, designed with matched compute, storage, and network scale, removes the estate-side imbalances that no amount of job tuning can.
FAQ
What MFU should a training cluster deliver?
Scale- and model-relative numbers, not one threshold: production systems demonstrate ~55% at twelve-thousand-GPU scale, and smaller clusters often land higher on well-fitting models while both beat the ~23% divergence cases — set your baseline from your own run history under a fixed convention, and treat published peaks as ceilings rather than promises.
Is MFU collapse an imbalance problem or an efficiency problem?
Both, and the timing split tells you which: if ranks finish unevenly and the slowest waits dominate, it is imbalance; if every rank is uniformly slower than its kernel-level potential, it is single-GPU efficiency; if the loss concentrates in bursts tied to one node or link, it is a straggler — research treats inter-GPU imbalance and single-GPU efficiency as separate measured factors, and so should your diagnosis.
When is low MFU a hardware problem rather than a configuration problem?
When the same configuration diverges across hardware: identical settings producing materially different MFU on equivalent nodes is the documented production-divergence signature, and the sequence is thermal check, health check, then drain — while uniform low MFU across all nodes points at configuration (parallelism shape, batch, kernels) long before it points at the metal.