Skip to content
HChuggingface.co·

Impactful scheduling for GPU clusters

AI summary

Ai2 manages thousands of NVIDIA H100, B200, and B300 GPUs in clusters of 88 to 1024 GPUs, supporting about 150 internal researchers across diverse AI domains like LLM/VLM training and robotics RL. To optimize scheduling and prioritize high-impact research while maintaining full occupancy, Ai2 introduced a “scheduling contract.” Workloads declare a minimum runtime for protected progress, after which they can be rebalanced. Alternatively, workloads with zero minimum runtime are subject to preemption but are not charged to any budget.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Oct 9, 2026, 15:20 UTC

IngestedOffset at this time: UTC+0Oct 9, 2026, 16:00 UTC

Published
Oct 9, 2026, 15:20
Ingested
Oct 9, 2026, 16:00
Source type
Official
Tier
First-party
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Building a cluster scheduler to prioritize high-impact research while maintaining full occupancy

On the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads. We think about this task as a pyramid of four metrics that build on each other.

The foundation is availability: how often the hardware is healthy and ready for work. Above this is occupancy: the fraction of available time assigned to a specific workload. Next is impact: how often the most valuable workloads are chosen to receive resources. The capstone of the pyramid is utilization: the fraction of GPU capacity used over the lifetime of a workload.

This post is about improving the impact of our scheduling decisions. We recently replaced a priority-based scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. As a result, we shifted the debate about how much GPU time each research project deserves from a case-by-case operational task to a transparent administrative budgeting process.

Overcommitting

At Ai2, we manage thousands of NVIDIA H100, B200, and B300 GPUs arranged in clusters ranging in size from 88 to 1024 GPUs. These clusters are built for large-scale distributed training of AI models, and they serve a group of about 150 internal researchers whose work covers a diverse set of AI domains, including the full model flow of LLM and VLM training, robotics reinforcement learning (RL) simulation, and post-training for scientific agentic use cases.

Like many labs, we have demand for GPU time that far exceeds supply. Based on submitted workloads, at any moment in time we have outstanding requests for 2-3x more GPUs than are available. One way to think about this is that every available GPU hour on our cluster has 2-3 different research workloads competing for it.

Historically, we used a priority-based scheduler, and we allowed workloads to opt out of preemptability. Each team had a limit on concurrent GPUs that could be used by workloads which were protected from preemption. Preemptible workloads could exceed that limit on idle GPUs. This strategy produced predictable pathologies. For example, we observed instances of GPU “squatting” where users would park no-op workloads they could connect to when the need arose. This occurred because researchers found they could not launch debugging workloads with low enough latency to tackle problems in real time. We also observed priority inflation, where eventually 100% of scheduled workloads used HIGH priority. This meant that lower priority levels were starved of GPU time altogether. Since preemptability was optional, we also found that our on-call engineers spent a majority of their ticket response time negotiating the organized shutdown of non-preemptable workloads running on hosts with known maintenance problems.

Tragedy of the commons

When these problems emerged, we were slow to identify their root causes. Our initial attempts to ensure the most important work received GPU time were focused on tighter control of how priorities were set, and, ultimately, working around the priority-based scheduler by explicitly assigning GPU monopolies to important projects. While we didn’t recognize it at first, we had built a perfect laboratory for observing the “tragedy of the commons.” Individuals were competing over a scarce, shared resource and, by seeking to maximize individual outcomes, achieving a non-optimal global result and abusing the underlying resource.

We were far from the first to observe this kind of interaction. Resource allocation is a fascinating research domain that mixes algorithm development, economics, and system management. A central problem is that users often know the value of their own jobs better than the organization does, but they may have incentives to hide that value or hold on to resources even when doing so hurts total performance. For example, in their 2011 paper introducing Dominant Resource Fairness, Ghodsi et al. recount an anecdote in which a search company provided dedicated machines to jobs only if their users could guarantee high utilization. They soon discovered “users would sprinkle their code with infinite loops to artificially inflate utilization levels.” The hardware changes, but the fundamental problems that make resource allocation complex persist.

Budgets not schedules

The classic solution to a tragedy of the commons is to privatize the shared resource—owners are incentivized to maximize the value of their property. When we assigned teams monopolies over sets of GPUs, we were already doing a version of this, but it was too coarse. It caused GPUs to sit idle due to the seasonality of research. Teams are ready to run experiments and training at different times, so assigning a monopoly would ensure that there would be times when no jobs were ready to execute, with another team left waiting for capacity.

We were manually solving a knapsack problem, trying to fit dynamically changing research needs into a static schedule. We wanted the ownership incentive, but we also wanted to maintain full occupancy of the GPUs.

We decided to iterate on the ownership model. Instead of issuing teams GPUs, we chose to allocate a portion of GPU time. Predicting demand into the future would require knowing the result of novel science experiments, so it cannot be forecast with precision. Priority across research efforts, however, is a question of strategy, and it can be more easily debated and decided in advance. Instead of trying to solve the scheduling puzzle, we enabled leadership to think like investors. Before the workloads exist, decide how to fund each research effort with GPU time based on their judgment of its likely impact. The scheduler could then use that information when prioritizing arriving workloads.

With this in mind, we devised a hierarchical system where managers could proportionally allocate GPU time to the projects and researchers they were responsible for. As the diagram below illustrates, this translates program strategy directly into a guaranteed share of GPU time. Project A1 knows it has a 35% claim on total capacity, regardless of how many other projects are queuing up elsewhere.

Parenthetical values represent the total cluster capacity assigned to a leaf project.

In this system, every request for GPU time must be funded by a budget, or it is not protected from preemption. In the old system, HIGH priority carried no cost and non-preemptibility allowed a team to fill their concurrent GPU limit indefinitely, so everyone used them. Now, nothing is free, so any trick to get GPU time draws from the benefiting user’s allocation. A squatting workload is spending team budget on nothing. Our strategy is to make gaming the scheduler more expensive than honestly engaging in the debate for a larger budget. We are constantly iterating on this budget review process, but the key requirements are that there are frequent opportunities for researchers to advocate for the time they need, and the decisions are made by managers with the most context on the tradeoffs in question. This means allocation decisions within a research project are made by a lead researcher, within a research program by a principal investigator, and across programs by a lead program manager, or by the CEO.

Fair-share

Paired with this GPU time budgeting tool, we built a hierarchical fair-share scheduler to manage actual occupancy of allocations throughout the program tree. The algorithm here is not new—hierarchical fair-share over a time window is part of a lineage that goes back to the Hadoop Fair Scheduler in 2009, and the same approach is in active use today in SLURM’s Fair Tree and YARN’s Fair Scheduler. What’s new for us are the inputs: the tree mirrors the research program structure, and the weights are budgets set by managers rather than static quotas.

The scheduler tracks occupancy over a sliding lookback window (we default to 7 days) and sorts workloads from under-utilized allocations above those from over-utilized allocations. This way, over a week-long time range, we can expect every group to receive their allocated GPU time as long as they are actively submitting workloads with sufficient demand.

“The new scheduler makes it feel like we have an extra 30% compute. In the old scheduler, if we had moments when we didn't need our full slot limit, that compute was basically lost. Now with the new scheduler, if that happens, we can later burst beyond our allocation limit and still see our jobs scheduled quickly and without preemption, essentially letting us reclaim that compute. Our workloads are often bursty, so this gave us a significant amount of compute back.” — Chris Clark

The scheduler distinguishes two kinds of occupancy. Allocated occupancy is time during which a workload is charged to a budget. This draws from the workload owner’s allocations, which affects the fair-share budget calculation, and these workloads are protected from preemption during their minimum runtime window. Unallocated occupancy is not charged to any budget, is unprotected from the outset, and may be preempted by any allocated request. This allows us to keep the GPUs fully occupied even when allocations don’t properly match demand and prevents teams from ever declining free GPU cycles.

The scheduling contract