Critical
78% of business-critical GPU capacity stands on one cooling loop
144 critical GPUs (atlas-70b-ft, rec-train-batch) share CDU loop A. A single cooling event stops them together — redundancy at the node level does not cover a shared loop.
144 of 184 critical GPUs on CDU loop A; est. $685 of training progress lost per failure event (half checkpoint interval + 90 min restart overhead).
Governed actionRedistribute 40 GPUs of the standing training work onto healthy free capacity in other cooling domains.
critical share on worst loop: 78% → 57%idle power added: +2.8 kWnode moves: 5
Critical
One power domain (PDU 1) carries 78% of critical capacity
144 critical GPUs sit behind PDU 1. A breaker or shelf event is a full stop for atlas-70b-ft.
144 critical GPUs behind one PDU; $567 progress at risk per event.
Governed actionSplit the run across power domains during the next scheduled restart; make power-domain spread a placement constraint.
single-PDU exposure: 144 → 72 GPUs
Warning
119 GPUs look free — 48 are safely usable
Free GPU count overstates usable capacity: fragmentation blocks 8-GPU gangs, a power-capped rack cannot energize, and lemon nodes fail the jobs placed on them.
raw free 119 = fragmented 15 + power-capped 32 + on lemon nodes 24 + safe 48.
Governed actionDefragment partial nodes, rebalance the rack power budget, and cordon lemon nodes — then re-admit queued gangs against real capacity.
schedulable 8-GPU gangs: 6 today → 7 after defrag
Warning
3 lemon nodes are quietly failing the jobs placed on them
n23, n35, n29 carry repeat failure history. Meta's production study: excluding nodes like these cut large-job failure rates from 14% to 4%.
n23: 41 ECC error events in 90d; 7 jobs died on this node in 90d; 12 driver XID events in 90d · n35: 28 ECC error events in 90d; 5 jobs died on this node in 90d; 6 driver XID events in 90d · n29: 19 ECC error events in 90d; 4 jobs died on this node in 90d; 5 driver XID events in 90d
Governed actionCordon, burn-in test, and re-admit only after a clean pass; placement should price node failure history until then.
free GPUs behind lemons: 24
Warning
Two control loops fought between 02:00 and 03:00
cluster-autoscaler and custom rebalancer (cron) acted on the same workloads in the same hour — 59 placement actions, each undoing the other's work. No shared objective arbitrates them.
cluster-autoscaler: 21, custom-rebalancer: 38, spot-replacer: 0 actions in hour 2.
Governed actionPut both loops under one migration budget with a shared cooldown; a governance weight decides which optimization wins when they disagree.
actions in collision hour: 59 → ~12
Info
57 allocated GPUs run under 15% utilization
Allocated is not utilized: these GPUs are held by workloads but doing almost no work — invisible in allocation dashboards, fully visible on the bill.
57 GPUs × $2.10/GPU-h × 720 h ≈ $86200/month.
Governed actionRight-size or time-box the holders; reclaim into the safe-usable pool with risk-priced consolidation (not blind bin-packing).
idle spend: $86200/mo