Ai2's GPU queue ended with every job marked high priority, so it moved to budgets
The Allen Institute for AI rebuilt its GPU scheduler around time budgets. Debug jobs that waited two hours now start in about 30 seconds.
OddBrief EditorialAI-assisted, human-reviewed
AIKey facts
- Who
- Ai2's AI infrastructure team, about 150 researchers, thousands of H100, B200 and B300 GPUs
- Problem
- priority inflation until 100% of jobs were HIGH, plus GPU squatting
- Result
- 98% of owed GPU hours delivered over 30 days; debug p90 wait 2 hours to 30 seconds
- Catch
- interactive sessions now become preemptible after 8 hours
The Allen Institute for AI (Ai2) replaced the priority system on its GPU clusters after, by its own account, eventually all scheduled workloads were marked high priority. In an October 9 post, the institute's AI infrastructure team says a budget-based scheduler has since delivered 98% of the GPU hours teams were owed, while cutting waits for short debugging jobs from 2 hours to 30 seconds.
The post is a rare look at how a large AI lab rations scarce compute among its own researchers, and at how people game the rules when nothing costs them anything.
A tragedy of the commons on H100s
Ai2 runs thousands of Nvidia H100, B200 and B300 GPUs in clusters of 88 to 1,024 GPUs, used by about 150 researchers. At any moment, the team says, submitted work asks for 2 to 3 times more GPUs than exist.
Under the old scheduler, jobs could pick a priority and could opt out of being preempted. Both were free, so both were overused. Priority inflated until every job was HIGH, which starved anything lower. Some researchers practiced what the team calls GPU squatting: parking idle placeholder jobs on GPUs so they could jump in quickly when they needed to debug, because real debug jobs waited too long.
Non-preemptible jobs also landed on on-call engineers, who spent most of their ticket time negotiating shutdowns of workloads sitting on machines that needed repair. The team says it was slow to see the root cause and first tried tighter priority rules and handing whole blocks of GPUs to important projects, which left hardware idle between experiments.
GPU time as a budget, not a slot
The new system hands out shares of GPU time rather than GPUs. Managers split a budget down a tree of programs, projects and researchers, and a fair-share scheduler compares each group's use with its share over a rolling 7-day window. Any job not funded by a budget can be preempted at any moment.
Each job must also declare a minimum runtime, capped at 8 hours, during which it cannot be interrupted. After that, the scheduler may pause and requeue it. Ai2 says the point is to make gaming the scheduler cost more than arguing honestly for a bigger budget. A squatting job now simply burns its own team's allocation.
What changed, and what got worse
Over a 30-day test, 13 of 15 team allocations received at least 95% of their owed hours, and the worst got 90%. Cluster occupancy stayed at 98% before and after. On the largest H100 cluster, the median wait in the queue dropped to 24 seconds, down from 5 minutes. Because jobs now drain on their own when a machine needs repair, repairs needing a human fell by 74%.
Researcher Chris Clark said the new scheduler "makes it feel like we have an extra 30% compute," because unused allocation can be reclaimed later.
Not everyone gained. Interactive sessions that researchers once held for up to a week now become preemptible after 8 hours, and losing one means rebuilding its state by hand. Ai2 says it is building a separate CPU-only cluster and restorable sessions to fix that, and is still checking whether time-slicing fragments capacity for its largest jobs.
Sources
- Impactful scheduling for GPU clustersAi2 (Hugging Face blog)primary source