Ai2 Rebuilt Its GPU Scheduler and Cut Wait Times by 87%

Ai2 replaced its priority-based GPU scheduler with time budgets and fair-share allocation, cutting median queue wait from 5 minutes to 24 seconds.

·
·
·
  • Ai2 replaced its priority-based GPU scheduler with time budgets, hierarchical fair-share, and a time-slicing contract.
  • Median queue wait on their largest H100 cluster fell from 5 minutes to 24 seconds.
  • Teams received 98% of owed GPU hours while cluster occupancy stayed at 98% under 2-3x oversubscription.
  • Jobs declare a minimum runtime (capped at 8 hours) during which they are preemption-protected.
  • Automated host draining via time-slicing cut human-in-the-loop repairs by 74%.
  • Full details in the Ai2 engineering blog.

Ai2 cuts GPU queue waits with time budgets

The Allen Institute for AI (Ai2) rebuilt the scheduler for its NVIDIA H100, B200, and B300 clusters after cost-free priority labels and long-lived protected jobs encouraged researchers to hoard capacity. The replacement combines GPU-hour budgets, hierarchical fair-share scheduling, and bounded protection from preemption.

On Ai2’s largest H100 cluster, median queue wait fell from 5 minutes to 24 seconds. The p90 wait, the threshold below which 90% of waits fall, declined from 2.8 hours to 1.8 hours. The design gives infrastructure teams a concrete way to allocate scarce accelerators while preserving high cluster occupancy.

Priority without a price

Ai2 operates thousands of GPUs across clusters containing 88 to 1,024 devices each. About 150 internal researchers use them for large language and vision-language models, robotics reinforcement learning simulations, and agent post-training. Queued demand often reaches two to three times available capacity.

The previous scheduler let workloads request protection from preemption, while team-level concurrency limits capped the number of protected jobs. Because high priority and protection carried no usage cost, the policy produced three operational failures:

  • GPU squatting: Researchers parked no-op workloads on accelerators so they could connect immediately, avoiding long waits for debugging sessions.
  • Priority inflation: Eventually, every scheduled workload used HIGH priority, starving jobs submitted at lower levels.
  • Maintenance gridlock: On-call engineers spent substantial time negotiating shutdowns of protected jobs on hosts that required repair.

Ai2 describes the incentives as a tragedy of the commons. Its account cites a similar example from the 2011 Dominant Resource Fairness paper, in which users added infinite loops to inflate utilization figures and retain dedicated machines.

The layer a scheduler can fix

Ai2 diagnoses cluster performance with a four-layer model that separates scheduling policy from hardware health and workload efficiency, with this redesign concentrating on the impact layer:

Availability
How often the hardware is healthy and ready to accept work.
Occupancy
How much available GPU time is assigned to workloads.
Impact
How often the scheduler selects the work the organization values most.
Utilization
How much GPU capacity a running workload consumes over its lifetime.
Ai2 pyramid showing availability, occupancy, impact, and utilization
Ai2’s model separates scheduler impact from fleet availability, occupancy, and workload utilization.

GPU-hours replace priority labels

Fixed GPU allotments left hardware idle because research demand arrives in bursts. Ai2 now allocates shares of total GPU time through a hierarchy of programs, projects, and researchers. These budgets provide weighted entitlements while allowing teams to borrow unused capacity.

For each allocation, the scheduler calculates a rolling seven-day ratio of actual occupancy to allocated GPU time. Jobs associated with underused allocations rank ahead of jobs whose teams have already consumed their share. A workload can continue beyond its requested minimum runtime while its allocation still outranks competing demand.

Because every charged workload consumes its team’s budget, a no-op squatting job reduces the GPU time available for later training. The accounting system aligns each scheduling decision with the allocation that benefits from it.

The algorithm follows a mature lineage that includes the 2009 Hadoop Fair Scheduler, YARN’s Fair Scheduler, and Slurm’s Fair Tree. Ai2 maps the hierarchy to its research organization and lets managers set weights as budgets instead of relying on static per-user quotas.

Eight hours of protection, then preemption

Large training jobs complicate fair-share scheduling because a week-long run can hold hundreds of GPUs after competing allocations become underserved. Ai2 addresses that problem with a workload contract: each job declares a minimum runtime and whether it can resume after eviction, with protected runtime capped at eight hours.

  1. The user submits the resource request, minimum runtime, and resumable flag.
  2. The scheduler ranks the job using its allocation’s actual-to-entitled occupancy ratio over the rolling lookback window.
  3. Once placed, the job consumes GPU time from that allocation.
  4. The scheduler protects it from preemption for the declared minimum runtime.
  5. After protection expires, the job may continue while its allocation outranks competitors. A resumable job may be evicted and returned to the queue.
  6. Completion releases the resources for another workload.

A job submitted with minimum_runtime=0 enters an unallocated backfill class. It does not consume a team budget, can use otherwise idle GPUs, and may be preempted immediately when allocated work arrives.

As protection windows expire, workloads can leave hosts marked for maintenance without engineers negotiating individual shutdowns. Ai2 reports a 74% reduction in repairs that require human coordination.

The rollout kept 98% occupancy

Before deployment, Ai2 built a simulator to test the seven-day lookback window, the maximum protected runtime, and other policy settings against historical traces and synthetic scenarios. A 30-day production rollout produced the following results:

Measure Production result
Budget delivery Teams received 98% of their allocated GPU-hours.
Allocation consistency 13 of 15 team allocations received at least 95% of their budgets; the lowest received 90%.
Cluster occupancy Held at 98% before and after the scheduler change.
Unallocated backfill Accounted for 18% of delivered GPU time.
Debug workload p90 wait Fell from 2 hours to 30 seconds. The simulator had modeled a reduction from 6 hours to 5 minutes.

Borrowing unused entitlements also lets bursty workloads temporarily exceed their assigned shares. One researcher described the effect as “an extra 30% compute” because capacity that previously sat idle became available between other teams’ runs.

Shorter queues carry trade-offs

Changed terminology created confusion during the transition because researchers still used words such as “priority” after their scheduling meaning had shifted. Written documentation proved insufficient, so the infrastructure team held live question-and-answer sessions built around actual scheduling decisions.

The eight-hour protection cap also disrupted researchers who kept interactive development sessions alive for a week while storing volatile state in memory. Preemption forced them to reconstruct that context manually. Ai2 plans to provide a separate CPU development cluster and restorable sessions for this workflow.

Minimum-runtime guarantees can also fragment capacity. Protected small jobs reduce the opportunities to reclaim enough GPUs simultaneously for a 512-GPU workload, so Ai2 continues to monitor placement delays for its largest runs.

The operating model travels

Ai2’s most reusable decision was to turn case-by-case operations disputes into GPU-hour allocations that leadership sets before jobs enter the queue. Managers distribute compute according to expected research value, while the scheduler measures delivery and enforces those decisions over time.

The policy can sit above Slurm, Kubernetes, or a custom scheduler because its core inputs are an organizational hierarchy, weighted budgets, metered occupancy, and explicit preemption rules. Implementing it requires several supporting systems:

  • Reliable accounting: Track GPU time by workload, researcher, project, and team over a defined lookback window.
  • Budget ownership: Give designated managers authority to assign and revise shares.
  • Preemption support: Encourage checkpointing and define bounded protection for resumable workloads.
  • Backfill capacity: Provide an immediately preemptible class that can absorb unused GPUs.
  • Rollout tooling: Simulate historical traces, explain ranking decisions, and provide a separate path for long-lived interactive development.

Ai2’s technical account covers further design trade-offs and planned work on GPU utilization. The organization also lists open infrastructure roles.

Trending
  • No trending articles

Comments

avatar

Next Reads