Ai2 has replaced a priority-based scheduler for its research clusters with a system built around GPU-time budgets, hierarchical fair sharing and mandatory time slicing. The change moves decisions about which projects deserve scarce compute from daily operations into an explicit budgeting process.

The institute manages thousands of Nvidia H100, B200 and B300 GPUs in clusters ranging from 88 to 1,024 accelerators. Roughly 150 researchers use them, and outstanding requests typically ask for two to three times the available capacity.

Under the old system, nearly every queued workload eventually carried a high-priority label. Some users kept idle jobs alive so they could obtain a machine quickly, while non-preemptible workloads made hardware maintenance a negotiation. Lower priorities could be starved even when their work was useful.

The replacement gives organizations and projects defined shares of GPU time, then lets the scheduler distribute capacity within that hierarchy. A time-slicing contract means work must tolerate interruption, while budgeting forces leadership to make trade-offs visibly rather than relying on inflated labels. Ai2 presents the design as a way to improve the impact of a busy cluster, not merely its occupancy; actual model throughput still depends on how efficiently each scheduled job uses its allocation.