GPU estate governance

Efficiency has a price. Now it has a price tag.

Every consolidation, curtailment, and cost optimization spends something the dashboard never shows: blast radius, thermal headroom, goodput at risk. Iverson GPU makes that price explicit — then turns the tradeoff dial for you, inside a policy you set.

Read-only. One Helm command. No workload changes.

Consolidate
energy · carbon · $ — one shared failure domain
Spread
resilience · headroom — idle watts everywhere

One governance weight spans the whole tradeoff — and the right setting moves with your workload. A weight tuned on one workload produced 52% SLA violations on another. That is why a controller turns this dial, not a human.

1 / 3 hrs

unexpected interruption rate while training Llama 3 on 16,384 GPUs — 419 in 54 days.

Meta, 2024

5–30%

of paid-for GPU capacity does useful work across measured production fleets.

CAST AI 23k-cluster telemetry · FinOps Foundation 2026

72 GPUs

share one NVLink backplane, cooling loop, and power shelf in a rack-scale system. One failure can take them all.

GB200 NVL72 architecture

11 hours

of global derivatives trading frozen by one shared cooling plant — behind fully redundant chillers.

CME / CyrusOne, Nov 2025

The problems nobody prices

Your dashboards show the symptoms. They rarely price the tradeoff.

Control-loop collision

Autoscalers, rebalancers, and optimizers act on the same pods with no shared objective. The result is churn, flapping, and the 2 a.m. migration storm nobody ordered.

What we do One governance weight arbitrates every loop. In the paper's benchmark, governed placement cut scheduling oscillation 32×.

Correlated-domain blindness

Schedulers see zones and labels. They do not see that eighteen nodes share a cooling loop, a power shelf, and a top-of-rack switch — until all of them fail together.

What we do Physical failure domains are first-class objects. Placement is priced by the blast radius it creates, before anything fails.

The usable-capacity illusion

The dashboard says 119 GPUs are free. Fragmentation, power caps, and quietly failing nodes mean 48 of them can actually take an 8-GPU job.

What we do A safe-usable-capacity ledger, decomposed and priced — the number capacity planning should have been using all along.

Recommendation debt

Optimization suggestions pile up unactioned because every one transfers un-quantified risk to the engineer who clicks apply.

What we do Every recommendation ships with its risk price, confidence, and rollback path — so accepting it becomes rational.

How it works

Read-only first. Trust is earned in rungs.

01

Observe

A read-only agent maps your estate: workloads, GPUs, and the physical domains they stand in. First exposure report within a day. No write access, ever.

02

Govern

Risk-priced recommendations land as pull requests against the policies you already run — consolidation, placement constraints, migration budgets. You approve; we show the receipts.

03

Self-tune

The controller turns the efficiency–resilience dial continuously, inside hard bounds you set. You choose the envelope; it finds the safe edge — and proves it did.

Built on peer-reviewed research

The governance controller is published, benchmarked, and reproducible.

Self-Tuning Governance of the Energy–Resilience Tradeoff in GPU Cluster Scheduling (IEMCON 2026) evaluates the controller on replayed Microsoft production traces and measured grid-carbon data. The simulator behind our sandbox reproduces the paper's reference implementation bit-for-bit — run the experiments yourself.

Honest footnote: results are trace-driven simulation, not customer-cluster measurements. The violation metric is conservative by construction; read paper numbers as paired comparisons, not production SLA figures.

52%

SLA violations when a statically tuned weight is transplanted onto a production trace — the case for self-tuning.

32×

less scheduling oscillation than greedy energy minimization under a load ramp.

5.5×

fewer jobs lost when a correlated high-centrality domain fails — exposure avoided before the failure, not recovered after it.