GPU Capacity Expansion Decision Rights: Who Decides, on What Evidence
Every GPU estate has a moment when someone asks for more — and in estates without assigned decision rights, that request begins the most reliable delay in infrastructure: the search for whoever can say yes. The demand forecast exists, the budget line exists; what is missing is the structure between them — who approves what size of growth, on whose evidence, at what speed. This page builds that structure.
Prerequisites: Assign Rights Before the Next Expansion
The rights map exists before the request, not during it: a RACI-style assignment with an accountable owner (the sponsor who approves material trade-offs and owns the outcome), a responsible platform function preparing the case, and consulted seats for finance and security — because governance playbooks are explicit that the accountable sponsor approves material trade-offs while functional roles advise, and the unassigned version of this map is the informal stall where every expansion invents its own approval path.
| Seat | Role in an expansion | What it contributes |
|---|---|---|
| Sponsor (accountable) | Owns the outcome; approves material trade-offs | The yes, and the accountability behind it |
| Platform (responsible) | Builds the case; executes the decision | Utilization data, options, technical judgment |
| Finance (consulted) | Validates the spend shape | Budget fit, commitment implications |
| Security/Compliance (consulted) | Validates the boundary | Tenancy, residency, control implications |
The four seats are the constant; the titles vary by org chart. What varies more — and costs more — is what happens without the map: expansion requests routed by personality and proximity, decisions re-litigated per incident, and the advisory finding that early alignment brings capacity online faster arriving as experience rather than practice.
Build the Tiers: Threshold, Evidence, Path
Tier the rights by threshold so the decision scales with the cost: routine growth (within existing budget and forecast bands) approves at platform level on utilization evidence; material expansion (new committed capacity) routes to leadership with forecast bands and downside scenarios; strategic commitments (multi-year, facility-scale) reach the executive sponsor with contract and scenario analysis — each tier consuming the evidence its size deserves, so a quarter-rack request never waits on the committee that a data-hall commitment must face, and vice versa.
- Set the thresholds: dollar or rack counts written down per tier — the numbers that route a request before opinion can.
- Match evidence to tier: utilization history for routine, forecast bands and downside scenarios for material, contract and scenario analysis for strategic — evidence scales because the reversal cost does.
- Pre-build the packs: each tier's evidence template exists before a request needs it, so preparation is assembly, not authorship.
- Announce the map: the structure published where requesters can read it — an unannounced governance model is just bureaucracy with better documentation.

The tiers also encode the consulting lesson in contract form: capacity, permits, and approvals belong in agreements as commitments with reporting, not as assumptions — the strategic tier is where those commitments get made deliberately.
Verify: Test the Paths Before You Need Them
Test the paths like any control: walk one routine and one material expansion through the structure on a real (small) request and measure the clock — decision latency per tier, evidence actually consumed versus demanded, and seats that could not be reached — because advisory practice is clear that aligning decisions early brings capacity online faster, and the test converts the rights map from a diagram into the operating speed your next expansion will actually experience.
The walkthrough also finds the seat that never responds — every organization has one — while the stakes are a drill rather than a stalled program, and it exposes the evidence packs nobody actually reads, which get trimmed now instead of burdening every future request. Governance that has never been walked is a hypothesis; the walkthrough is the experiment, and the next real expansion is a terrible place to run it.
FAQ
What approval thresholds should a GPU estate set?
Three tiers with numbers attached: routine growth inside existing budget and forecast bands (platform approves on utilization evidence), material expansion adding committed capacity (leadership approves on forecast and downside scenarios), and strategic commitments (executive sponsor on contract and scenario analysis) — the specific dollar or rack counts are estate-specific, but the tiers exist so the decision scales with the cost instead of every request negotiating its own path.
Who should own the GPU capacity decision — platform, finance, or the business?
A sponsor from the business owns it, with platform preparing the case and finance and security consulted: the decision-rights pattern puts one accountable owner on the outcome while the functional seats advise, because expansions fail differently under single-function ownership — platform-owned grows without cost discipline, finance-owned grows without technical judgment, and the sponsor model forces both into one decision.
How should urgent GPU expansions be handled without breaking governance?
With a designed emergency path, not an improvised one: a named emergency approver, a reduced evidence pack (the need, the cost, the rollback), immediate routing to the tier that normally owns it, and mandatory post-hoc review at the next governance cycle — the same break-glass pattern release processes use, so urgency changes the speed of the decision, never its accountability. Dedicated providers with predictable flat-rate commitments, such as OneSource Cloud, shorten the strategic tier for this reason: the commitment decision is simpler to make fast when the terms are simple to read.