Ai2 replaces its GPU priority scheduler with time budgets and fair-share

Ai2's AI Infrastructure team, which provides the institute's GPU capacity for large distributed training workloads, describes how it replaced a priority-based scheduler with a system of GPU time budgets, hierarchical fair-share allocation and a time-slicing 'scheduling contract'. The team thinks about GPU provisioning as a pyramid of four metrics that build on each other: availability (is the hardware healthy), occupancy (the fraction of available time assigned to a workload), impact (how often the most valuable workloads are chosen to receive resources) and utilization (the fraction of GPU capacity used over a workload's lifetime). The post is about impact.
The scale: Ai2 manages thousands of NVIDIA H100, B200 and B300 GPUs, arranged in clusters from 88 to 1024 GPUs each, serving about 150 internal researchers who work on LLM and VLM training, robotics reinforcement learning simulation and post-training for scientific agentic use cases. Demand far exceeds supply: based on submitted workloads, outstanding requests are for 2-3x more GPUs than are available, so every GPU hour has 2-3 workloads competing for it.
The old setup used a priority-based scheduler in which workloads could opt out of preemptability. Each team had a cap on concurrent GPUs for protected workloads, and preemptible workloads could exceed it on idle GPUs. The team saw three pathologies. First, GPU 'squatting': users parked no-op workloads they could connect to when the need arose, because they could not launch debugging workloads with low enough latency to tackle problems in real time. Second, priority inflation: eventually 100% of scheduled workloads used HIGH priority, which starved the lower levels of GPU time entirely. Third, on-call toil: because preemptability was optional, engineers spent a majority of their ticket response time negotiating the shutdown of non-preemptable workloads on hosts with known maintenance problems.
The team was slow to find the root cause. Early fixes were tighter control over how priorities were set and, ultimately, working around the scheduler by assigning GPU monopolies to important projects. Monopolies were too coarse: GPUs sat idle because teams are ready to run experiments at different times. The authors came to see the situation as a tragedy of the commons and cite the 2011 Dominant Resource Fairness paper by Ghodsi et al., which recounts a search company that gave dedicated machines only to jobs whose users guaranteed high utilization, and found that users would sprinkle their code with infinite loops to artificially inflate utilization levels.
The new design allocates a portion of GPU time rather than GPUs. Demand cannot be forecast, since it depends on the results of novel experiments, but priority across research efforts is a strategic question that can be debated in advance. So leadership 'thinks like investors': before workloads exist, managers fund each research effort with GPU time based on their judgment of its likely impact. Budgets form a hierarchy. Allocation decisions within a project are made by a lead researcher, within a program by a principal investigator, and across programs by a lead program manager or the CEO. In the post's illustrative diagram, Project A1 has a 35% claim on total capacity regardless of how many other projects are queuing. Every request for GPU time must be funded by a budget or it is not protected from preemption, so nothing is free: a squatting workload is spending team budget on nothing. The stated aim is to make gaming the scheduler more expensive than honestly arguing for a larger budget.
A hierarchical fair-share scheduler enforces the budgets. The authors say the algorithm is not new, tracing it to the Hadoop Fair Scheduler in 2009 and noting that SLURM's Fair Tree and YARN's Fair Scheduler use the same approach. What is new for Ai2 is the inputs: the tree mirrors the research program structure and the weights are budgets set by managers rather than static quotas. The scheduler tracks occupancy over a sliding lookback window, defaulting to 7 days, and sorts workloads from under-utilized allocations above those from over-utilized ones. Over a week, every group should receive its allocated time as long as it keeps submitting workloads with enough demand. Occupancy comes in two kinds. Allocated occupancy is charged to a budget and protected from preemption during the minimum runtime. Unallocated occupancy is charged to nothing, is unprotected from the start and can be preempted by any allocated request, which keeps GPUs fully occupied when allocations do not match demand.
The 'scheduling contract' deals with very long training jobs, which can run for hours, days or weeks and would otherwise hold their GPUs for a week or more. In exchange for cluster access, a workload must declare its minimum runtime, the shortest occupancy needed to make meaningful progress. It is protected from preemption for that time; afterwards the scheduler may rebalance and automatically requeue resumable workloads. A user can set minimum runtime to zero, which makes the GPU time unallocated: always preemptible, but free of charge. The lifecycle is: submit with a minimum runtime and a resumable flag; get scheduled by fair-share, weighted by the ratio of actual occupancy to allocated time in the window; run for the minimum runtime, charged to the allocations; keep running while the allocations still prioritize the workload; possibly be preempted and requeued; finally complete and release resources.
Time-slicing also lets unhealthy hosts drain their workloads as minimum runtimes expire, so repairs can be automated. The authors say this was more important than they expected and that it reduced repairs requiring a human in the loop by 74%, a large saving in on-call toil. A user, Chris Clark, says the new scheduler 'makes it feel like we have an extra 30% compute' because bursty workloads can later exceed their allocation and still be scheduled quickly without preemption. The post then moves to a section on simulations, noting that policy changes can have unintended consequences and that the problem is zero-sum.
Key facts
- Ai2 replaced a priority-based GPU scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing 'scheduling contract' built on declared minimum runtimes.
- The old scheduler produced GPU squatting, priority inflation (eventually 100% of workloads used HIGH priority) and on-call engineers spending most of their ticket time negotiating shutdowns of non-preemptable jobs.
- Budgets are set by managers in a hierarchy (lead researcher, principal investigator, lead program manager or CEO) and enforced by a fair-share scheduler tracking occupancy over a sliding window that defaults to 7 days.
- Workloads with a minimum runtime of zero are unallocated: free of charge but always preemptible, which keeps GPUs occupied.
- The authors report a 74% reduction in repairs requiring a human in the loop, because unhealthy hosts can drain workloads as minimum runtimes expire.
Why it matters
Many labs face demand for GPU time well above supply; at Ai2 outstanding requests run 2-3x above available GPUs. The post shows what happens when a priority queue is the only tool for sharing that scarcity: everyone ends up at HIGH priority, and people park idle jobs to hold their place. The authors' answer moves the argument about who deserves GPU time from a case-by-case operational task to a transparent administrative budgeting process. The fair-share algorithm is not new; the contribution is wiring it to the research program structure with manager-set budgets as weights, and pairing it with minimum runtimes.
Who it affects
Directly, the roughly 150 internal Ai2 researchers whose workloads run on the clusters, the on-call engineers who handle host maintenance, and the managers, principal investigators and program leads (up to the CEO) who now set GPU time budgets. Anyone running a shared GPU cluster under heavy contention can read it as a worked example of replacing priorities with budgets, though the post describes Ai2's own setup rather than a product.
How to use it
There is nothing to install; the post is a design description. The reusable pattern is: fund every request for GPU time from a budget, and leave unfunded work unprotected; mirror the budget tree on how research is organized; track occupancy over a sliding window (Ai2 defaults to 7 days) and order the queue by under- or over-use of each allocation; require each workload to declare a minimum runtime and say whether it is resumable; offer a zero-runtime, free, always-preemptible tier for opportunistic work; and let hosts needing repair drain as minimum runtimes expire.
How solid is it
This is a first-party engineering account from the team that built and runs the system, so the description of the design and the old problems is direct. The numbers are thin on context. The 74% reduction in human-in-the-loop repairs comes with no baseline count and no timescale. The 30% figure is a user's subjective statement that the scheduler 'feels like' extra compute, not a measured gain. The 35% for Project A1 is an illustrative diagram example, not a real allocation. The 100% HIGH-priority figure describes the old system's eventual state. The text reviewed ends at the start of the 'Simulations' section, so any simulation results and conclusions are not covered here.
Risks and caveats
The authors themselves note that scheduling policy changes can have unintended consequences and that the problem is zero-sum: time given to one researcher is taken from another. Budgets rest on managers' judgment of likely impact, since demand cannot be forecast, and the team says it is still iterating on the budget review process. The guarantee that each group receives its allocated time holds over a week-long range and only for groups that keep submitting workloads with enough demand. Protection from preemption lasts only for the declared minimum runtime, and automatic requeueing applies to resumable workloads, so jobs that cannot be interrupted fit the contract poorly.
“The new scheduler makes it feel like we have an extra 30% compute.”
— Chris Clark, quoted in the Ai2 post