Hardware

Ai2 deploys budget scheduler across large GPU clusters

The Allen Institute for AI overhauled its cluster resource management with a budget-based, fair-share scheduler, drastically reducing queue delays while preserving full hardware occupancy.

Hugging Face Blog1 day agoHardware
Image: Hugging Face Blog

To eliminate bottlenecking across thousands of NVIDIA H100, B200, and B300 GPUs, the Allen Institute for AI replaced its legacy priority system with a budgeted, hierarchical fair-share scheduler. Operating clusters sized from 88 to 1024 GPUs, the institute supports roughly 150 internal researchers whose outstanding requests routinely exceed physical compute capacity by 2-3x. The updated approach replaces manual, case-by-case compute disputes with an administrative process where leadership pre-allocates guaranteed shares of GPU time to specific research programs.

Under the new design, workloads draw against assigned team budgets and agree to a scheduling contract. Users define a minimum runtime—capped at a maximum of 8 hours—during which their jobs cannot be interrupted. Once that threshold passes, or if jobs run without allocated funding, jobs become preemptible to let other funded tasks run. A 7-day sliding lookback window tracks usage to prioritize under-utilized groups. Unallocated GPU capacity allows teams to burst beyond baseline limits without losing idle cycles. Researcher Chris Clark noted that the setup "makes it feel like we have an extra 30% compute."

During a 30-day test period, overall cluster occupancy remained stable at 98%, with 18% of delivered GPU hours stemming from unallocated cycles. The system delivered 98% of total owed hours to teams, with 13 of 15 team allocations receiving at least 95% of their budgeted compute and the lowest receiving 90%. Furthermore, automation linked to the runtime boundaries decreased human-in-the-loop maintenance repair tickets by 74%.

The system also drastically reduced delay times across cluster tiers. For small debug workloads needing 15 minutes or less, 90th-percentile queue wait times plummeted from 2 hours to 30 seconds, outperforming simulator predictions that anticipated a drop from 6 hours to 5 minutes. On the institute's largest H100 cluster, median queue wait times fell from 5 minutes to 24 seconds, while 90th-percentile wait times dropped from 2.8 hours to 1.8 hours.

This is our own summary of reporting by Hugging Face Blog

More in Hardware