GPU estate governance
Every consolidation, curtailment, and cost optimization spends something the dashboard never shows: blast radius, thermal headroom, goodput at risk. Iverson GPU makes that price explicit — then turns the tradeoff dial for you, inside a policy you set.
Read-only. One Helm command. No workload changes.
One governance weight spans the whole tradeoff — and the right setting moves with your workload. A weight tuned on one workload produced 52% SLA violations on another. That is why a controller turns this dial, not a human.
unexpected interruption rate while training Llama 3 on 16,384 GPUs — 419 in 54 days.
Meta, 2024
of paid-for GPU capacity does useful work across measured production fleets.
CAST AI 23k-cluster telemetry · FinOps Foundation 2026
share one NVLink backplane, cooling loop, and power shelf in a rack-scale system. One failure can take them all.
GB200 NVL72 architecture
of global derivatives trading frozen by one shared cooling plant — behind fully redundant chillers.
CME / CyrusOne, Nov 2025
The problems nobody prices
Autoscalers, rebalancers, and optimizers act on the same pods with no shared objective. The result is churn, flapping, and the 2 a.m. migration storm nobody ordered.
What we do One governance weight arbitrates every loop. In the paper's benchmark, governed placement cut scheduling oscillation 32×.
Schedulers see zones and labels. They do not see that eighteen nodes share a cooling loop, a power shelf, and a top-of-rack switch — until all of them fail together.
What we do Physical failure domains are first-class objects. Placement is priced by the blast radius it creates, before anything fails.
The dashboard says 119 GPUs are free. Fragmentation, power caps, and quietly failing nodes mean 48 of them can actually take an 8-GPU job.
What we do A safe-usable-capacity ledger, decomposed and priced — the number capacity planning should have been using all along.
Optimization suggestions pile up unactioned because every one transfers un-quantified risk to the engineer who clicks apply.
What we do Every recommendation ships with its risk price, confidence, and rollback path — so accepting it becomes rational.
How it works
A read-only agent maps your estate: workloads, GPUs, and the physical domains they stand in. First exposure report within a day. No write access, ever.
Risk-priced recommendations land as pull requests against the policies you already run — consolidation, placement constraints, migration budgets. You approve; we show the receipts.
The controller turns the efficiency–resilience dial continuously, inside hard bounds you set. You choose the envelope; it finds the safe edge — and proves it did.
Built on peer-reviewed research
Self-Tuning Governance of the Energy–Resilience Tradeoff in GPU Cluster Scheduling (IEMCON 2026) evaluates the controller on replayed Microsoft production traces and measured grid-carbon data. The simulator behind our sandbox reproduces the paper's reference implementation bit-for-bit — run the experiments yourself.
Honest footnote: results are trace-driven simulation, not customer-cluster measurements. The violation metric is conservative by construction; read paper numbers as paired comparisons, not production SLA figures.
SLA violations when a statically tuned weight is transplanted onto a production trace — the case for self-tuning.
less scheduling oscillation than greedy energy minimization under a load ramp.
fewer jobs lost when a correlated high-centrality domain fails — exposure avoided before the failure, not recovered after it.