Sakhanda Wire
NVDA $229.43 -0.46% MSFT $535.03 +2.38% GOOGL $354.00 +1.64% META $725.17 +0.59% AMZN $261.36 +2.87%
← Back to the news

Impactful scheduling for GPU clusters

Impactful scheduling for GPU clusters
Enterprise Article
Published October 9, 2026
Building a cluster scheduler to prioritize high-impact research while maintaining full occupancy

Impactful Scheduling for GPU Clusters - Google Docs-image-1 (1)

On the AI Infrastructure team at Ai2, we’re responsible for providing the institute’s GPU compute capacity, specifically targeting large, distributed training workloads. We think about this task as a pyramid of four metrics that build on each other.

The foundation is availability: how often the hardware is healthy and ready for work. Above this is occupancy: the fraction of available time assigned to a specific workload. Next is impact: how often the most valuable workloads are chosen to receive resources. The capstone of the pyramid is utilization: the fraction of GPU capacity used over the lifetime of a workload.

Pyramid of GPU compute metrics, from availability at the base through occupancy and impact to utilization at the top.

This post is about improving the impact of our scheduling decisions. We recently replaced a priority-based scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. As a result, we shifted the debate about how much GPU time each research project deserves from a case-by-case operational task to a transparent administrative budgeting process.

Overcommitting

At Ai2, we manage thousands of NVIDIA H100, B200, and B300 GPUs arranged in clusters ranging in size from 88 to 1024 GPUs. These clusters are built for large-scale distributed training of AI models, and they serve a group of about 150 internal researchers whose work covers a diverse set of AI domains, including the full model flow of LLM and VLM training, robotics reinforcement learning (RL) simulation, and post-training for scientific agentic use cases.

Like many labs, we have demand for GPU time that far exceeds supply. Based on submitted workloads, at any moment in time we have outstanding requests for 2-3x more GPUs than are available. One way to think about this is that every available GPU hour on our cluster has 2-3 different research workloads competing for it. 

Historically, we used a priority-based scheduler, and we allowed workloads to opt out of preemptability. Each team had a limit on concurrent GPUs that could be used by workloads which were protected from preemption. Preemptible workloads could exceed that limit on idle GPUs. This strategy produced predictable pathologies. For example, we observed instances of GPU “squatting” where users would park no-op workloads they could connect to when the need arose. This occurred because researchers found they could not launch debugging workloads with low enough latency to tackle problems in real time. We also observed priority inflation, where eventually 100% of scheduled workloads used HIGH priority. This meant that lower priority levels were starved of GPU time altogether. Since preemptability was optional, we also found that our on-call engineers spent a majority of their ticket response time negotiating the organized shutdown of non-preemptable workloads running on hosts with known maintenance problems.

Tragedy of the commons

When these problems emerged, we were slow to identify their root causes. Our initial attempts to ensure the most important work received GPU time were focused on tighter control of how priorities were set, and, ultimately, working around the priority-based scheduler by explicitly assigning GPU monopolies to important projects. While we didn’t recognize it at first, we had built a perfect laboratory for observing the “tragedy of the commons.” Individuals were competing over a scarce, shared resource and, by seeking to maximize individual outcomes, achieving a non-optimal global result and abusing the underlying resource. 

We were far from the first to observe this kind of interaction. Resource allocation is a fascinating research domain that mixes algorithm development, economics, and system management. A central problem is that users often know the value of their own jobs better than the organization does, but they may have incentives to hide that value or hold on to resources even when doing so hurts total performance. For example, in their 2011 paper introducing Dominant Resource Fairness, Ghodsi et al. recount an anecdote in which a search company provided dedicated machines to jobs only if their users could guarantee high utilization. They soon discovered “users would sprinkle their code with infinite loops to artificially inflate utilization levels.” The hardware changes, but the fundamental problems that make resource allocation complex persist.

Budgets not schedules

The classic solution to a tragedy of the commons is to privatize the shared resource—owners are incentivized to maximize the value of their property. When we assigned teams monopolies over sets of GPUs, we were already doing a version of this, but it was too coarse. It caused GPUs to sit idle due to the seasonality of research. Teams are ready to run experiments and training at different times, so assigning a monopoly would ensure that there would be times when no jobs were ready to execute, with another team left waiting for capacity. 

We were manually solving a knapsack problem, trying to fit dynamically changing research needs into a static schedule. We wanted the ownership incentive, but we also wanted to maintain full occupancy of the GPUs.

We decided to iterate on the ownership model. Instead of issuing teams GPUs, we chose to allocate a portion of GPU time. Predicting demand into the future would require knowing the result of novel science experiments, so it cannot be forecast with precision. Priority across research efforts, however, is a question of strategy, and it can be more easily debated and decided in advance. Instead of trying to solve the scheduling puzzle, we enabled leadership to think like investors. Before the workloads exist, decide how to fund each research effort with GPU time based on their judgment of its likely impact. The scheduler could then use that information when prioritizing arriving workloads.

With this in mind, we devised a hierarchical system where managers could proportionally allocate GPU time to the projects and researchers they were responsible for. As the diagram below illustrates, this translates program strategy directly into a guaranteed share of GPU time. Project A1 knows it has a 35% claim on total capacity, regardless of how many other projects are queuing up elsewhere.

Hierarchical allocation of GPU time across research programs and projects; project A1 has a 35% share of total capacity.

Parenthetical values represent the total cluster capacity assigned to a leaf project.

In this system, every request for GPU time must be funded by a budget, or it is not protected from preemption. In the old system, HIGH priority carried no cost and non-preemptibility allowed a team to fill their concurrent GPU limit indefinitely, so everyone used them. Now, nothing is free, so any trick to get GPU time draws from the benefiting user’s allocation. A squatting workload is spending team budget on nothing. Our strategy is to make gaming the scheduler more expensive than honestly engaging in the debate for a larger budget. We are constantly iterating on this budget review process, but the key requirements are that there are frequent opportunities for researchers to advocate for the time they need, and the decisions are made by managers with the most context on the tradeoffs in question. This means allocation decisions within a research project are made by a lead researcher, within a research program by a principal investigator, and across programs by a lead program manager, or by the CEO.

Fair-share

Paired with this GPU time budgeting tool, we built a hierarchical fair-share scheduler to manage actual occupancy of allocations throughout the program tree. The algorithm here is not new—hierarchical fair-share over a time window is part of a lineage that goes back to the Hadoop Fair Scheduler in 2009, and the same approach is in active use today in SLURM’s Fair Tree and YARN’s Fair Scheduler. What’s new for us are the inputs: the tree mirrors the research program structure, and the weights are budgets set by managers rather than static quotas.

The scheduler tracks occupancy over a sliding lookback window (we default to 7 days) and sorts workloads from under-utilized allocations above those from over-utilized allocations. This way, over a week-long time range, we can expect every group to receive their allocated GPU time as long as they are actively submitting workloads with sufficient demand. 

“The new scheduler makes it feel like we have an extra 30% compute. In the old scheduler, if we had moments when we didn't need our full slot limit, that compute was basically lost. Now with the new scheduler, if that happens, we can later burst beyond our allocation limit and still see our jobs scheduled quickly and without preemption, essentially letting us reclaim that compute. Our workloads are often bursty, so this gave us a significant amount of compute back.” — Chris Clark

The scheduler distinguishes two kinds of occupancy. Allocated occupancy is time during which a workload is charged to a budget. This draws from the workload owner’s allocations, which affects the fair-share budget calculation, and these workloads are protected from preemption during their minimum runtime window. Unallocated occupancy is not charged to any budget, is unprotected from the outset, and may be preempted by any allocated request. This allows us to keep the GPUs fully occupied even when allocations don’t properly match demand and prevents teams from ever declining free GPU cycles.

The scheduling contract

An additional feature of distributed training that makes fair resource allocation difficult is that workloads can run for a very long time. Training jobs regularly run for hours, days, and sometimes even weeks. Once scheduled, a workload could remain on its assigned GPUs for a week or more, providing no opportunity for others to receive their budgeted time. This is the system property that made GPU squatting possible. It’s also what forced on-call engineers to negotiate with long-running job owners to address ongoing maintenance issues.

To address these problems, we introduced a “scheduling contract.” In exchange for access to the cluster, a workload must declare its minimum runtime, or the shortest amount of occupancy required to make meaningful progress. During this time, a workload is protected from preemption. This gives the researcher a guarantee of progress, while giving the scheduler the right to rebalance once that progress is banked, automatically re-queueing resumable workloads. Alternatively, a user can set minimum runtime to zero, which indicates that the GPU time should be unallocated. These workloads are always subject to preemption, but they’re also free in the sense that they are not charged to any budget.

The workload lifecycle follows this pattern:

  1. The workload is submitted with a minimum runtime and indicates whether or not it is resumable.
  2. The workload is scheduled according to the fair-share algorithm, weighted by a ratio of actual occupancy to allocated time in the lookback window.
  3. The workload runs for its minimum runtime, which is charged to its allocations.
  4. The workload may continue running as long as the associated allocations continue to prioritize it over others. This time is also charged to its allocations.
  5. It may be preempted and requeued, which returns to step 2.
  6. The workload completes, releasing its claim on any resources.

Together, these agreements add time-slicing to our scheduler. Running workloads may be removed and requeued automatically, allowing fair-share to converge and disincentivizing squatting. They also let unhealthy hosts drain their workloads as they reach their minimum runtimes, so repair activities can be fully automated. This last point was more important than we realized when planning this work. It reduced repairs requiring a human-in-the-loop by 74%, which was a massive savings in on-call toil.

Simulations

We know that scheduling policy changes can have unintended consequences. The zero-sum nature of the problem means that giving time to one researcher means taking away from another. Users who lose this exchange tend to look for new workarounds. Before rolling out the budget-based system, we wanted a fast way to predict where those longer wait times might arise and to test configuration knobs like the length of the lookback window or the maximum value to allow for minimum runtime (we chose 8 hours).

We built a small simulation environment that takes a set of workloads and their submission schedule as input and allows the scheduler to make preemption and GPU assignment decisions. With the knowledge of each workload’s requested number of GPUs and total runtime, the simulator could jump ahead to schedulable moments and provide analysis of queue wait times, preemption events, and distribution of GPU time across projects for many simulated days in a few seconds. We ran the simulator against both historical submission data and constructed scenarios we wanted to better understand.

One hypothesis we wanted to test involved “debug workloads.” These jobs require a small number of GPUs and a minimum runtime of 15 minutes or less, which is enough for the user to see whether a job launches successfully or crashes early due to a bug or misconfiguration. We wanted to know whether these jobs would see a shorter queue wait time than larger training workloads, which often need many GPUs and hours of runtime to make meaningful progress. Intuitively, these smaller jobs should rise to the top of the queue, since a small job can fit more places than a large one. But the precise queue latency was important. A short wait of a minute or two would unlock a new development practice, but a ten-minute wait becomes infeasible.

Our simulations required hand-built test case data, because our historical record did not contain a high enough volume of these debug-like workloads. Our results supported the hypothesis, showing p90 debug workload wait times fall from about 6 hours to only 5 minutes.

Simulator timelines comparing GPU occupancy under the baseline scheduler and the new allocations scheduler.

Smaller-scale simulator visualization of the baseline (left) and the new “allocations” scheduler (right). Each row is one GPU; each bar is a job, colored by parent workload with one color hue per team; hatching marks time where a job is interruptible, and a red edge marks a preemption. In the baseline, long urgent jobs are never interrupted, and fewer preemptions occur on lower-priority-level work. The new scheduler has a larger mix of colors on each GPU, illustrating occupancy rotation across teams.

Results

With simulation results in hand, we began a cluster-by-cluster rollout at the end of July. The results we care about are whether the workloads we chose to fund received their time, whether the new system maintained full occupancy, and whether researchers could reason about the scheduler to make informed decisions. 

Since rollout, we have observed users and teams consistently receiving their allocated GPU time. We count the time owed to a team as its allocation capped hour by hour at its actual demand. Over the 30-day test period, teams were delivered 98% of the GPU hours they were owed, and 13 of 15 team allocations received 95% or more with the worst case receiving 90%. Occupancy on the cluster held steady at 98% before and after the change, with demand exceeding capacity by 2-3x in both periods. 18% of delivered GPU time was unallocated, which is how we maintained high occupancy during periods when funded use cases were not ready to run.

Our simulator results proved to be directionally accurate with real outcomes overperforming our predictions. Debug workload p90 queue wait time fell from 2 hours to 30 seconds under the new scheduler, against a simulated prediction of 6 hours to 5 minutes from hand-crafted test scenarios. It’s worth noting that the smaller sample size of debug workloads in the baseline meant there was higher variance in those measurements. Queue latency in general improved as a side effect of time-slicing: on our largest H100 cluster, median queue wait time fell from 5 minutes to 24 seconds, and p90 wait time fell by about a third (from 2.8 hours to 1.8 hours).

Compared against the three problems we set out to solve:

  1. Squatting: Short debug workloads start in under a minute, reducing the value of squatting. The cost of this behavior charges the squatter’s budget, which prevents them from receiving time when they really need it.

  2. Priority inflation: We still allow workloads to declare priority, but it only impacts sorting within a team. Managers are incentivized to monitor priority across the group to optimize the use of their budgets.

  3. On-call toil: Unhealthy hosts drain automatically as workloads reach minimum runtime. Repairs requiring a human-in-the-loop fell by 74%.

Challenges

The learning curve was steeper than we had assumed. We rolled out the change incrementally, so in the early days researchers experienced varied behavior depending on which cluster they targeted. Furthermore, our interfaces retained some old terminology (like workload priority) whose meanings had changed. Documentation alone did not resolve the confusion. What did work was conducting live explanatory sessions, providing a forum for researchers to ask questions and for the engineering team to provide deeper descriptions of both how and why the scheduler made its prioritization decisions using real examples.

This was a key moment because it marked a pivot from an early period of frustration and folk theories to the current mode, where research groups communicate more often and more broadly about the GPU needs of their experiments. Researchers are now engaging in the budgeting discussion with clearer knowledge of the tradeoffs being made to accommodate any new request.

In addition to in-person sessions, we introduced new visualizations post-launch to give users a better sense of how closely their assigned GPU time is tracking their expected allocations and directly exposing the metric used for sorting the workload queue. This provided a simple place to look when a workload was preempted to understand why. These visualizations also helped budget owners, who could see how GPU time was being used across various projects under their management. 

Allocation usage over time, showing how delivered GPU time tracks expected allocations.

Example visualization of allocation usage over time.

Not every use case improved. In addition to distributed training, our researchers launch interactive sessions where they perform data analysis and test training code while they write it. In the old system, a researcher could hold such a session for up to a week. With time-slicing, they were subject to the 8-hour cap on protected runtime, after which a session becomes preemptible if it exceeds its allocation. We did not appreciate the extent to which researchers were dependent on the volatile state of these sessions. Being preempted meant waiting to secure a new session and also rebuilding their state by hand. After surveying researchers to understand the breadth of this problem, we created two new roadmap projects. We are investing in a CPU-only cluster next to our on-prem storage for dev sessions focused on data prep tasks. This will preserve our training cluster capacity for workloads that truly need it. Additionally, we plan to build restorable sessions for these CPU-only workloads. This will allow us to continue to preempt workloads at the end of their minimum runtime for maintenance or time-slicing, while also being able to restore the session elsewhere without the researcher rebuilding it. We can retain the operational and scheduling benefits of this new system while also improving the user experience.

We remain on the lookout for emerging issues. One potential problem we’re investigating is capacity fragmentation, which may result in queue wait times increasing for the largest workloads. Our intuition is that minimum runtime protection is being applied to the types of jobs that used to rely on preemptible mechanisms to exceed their team concurrent GPU limit. Previously, those jobs could be interrupted at any time, potentially wasting that time, but also making it easier to schedule large jobs. Now the scheduler may have fewer opportunities to interrupt many jobs at once to place a large pending workload. We’re currently using our simulator tools to reproduce this problem while also measuring the ground truth in production.

The future

As we look beyond the scheduling work described here, we are aiming at the capstone of the pyramid: utilization. We need to ensure that bootstrapping, checkpointing, and the training applications themselves are all done as efficiently as possible, maximizing the value of the scheduled time each workload receives.

If you’d like to tackle challenges like these in close collaboration with researchers, we encourage you to explore open engineering roles at Ai2.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Originally published by Hugging Face on

Read the original on Hugging Face ↗

Text and images are the property of Hugging Face and are reproduced here with attribution and a link to the original publication.

← Back to the news

More stories

All the latest news