In the world of AI infrastructure, "keeping GPUs busy" is often cited as the gold standard. However, new research from Dharma AI demonstrates that simple occupancy is not enough. By moving from a traditional FIFO (First-In-First-Out) scheduler to a constraint-aware allocator, the team achieved an increase in utilization by 33 percentage points and a jump in priority-weighted output by as much as 105%.
The Hidden Cost of FIFO
The problem with FIFO systems is twofold: reservations and ordering. In traditional models, real-time inference requires fixed resource reservations to meet peak demand, leaving GPUs idle during off-peak hours. Dharma AI likens these GPUs to "grounded aircraft"—earning nothing while remaining unavailable for other tasks.
The new allocator treats demand as a curve rather than a ceiling. It allows batch tasks, such as model training, to occupy the troughs of real-time demand, ensuring that critical services have the resources they need exactly when traffic spikes.
Five Constraints for Legal Allocation
To ensure a valid and high-scoring allocation, the system adheres to five strict rules:
- A GPU serves at most one job per timestep.
- Batch-like jobs occupy contiguous blocks of GPUs, sized to a power of two.
- Real-time jobs have a hard cap on how many GPUs they may swap between steps.
- A job that has started cannot be interrupted (no preemption).
- Every job respects its specific demand range.
Forecasting Over Reaction
The system's success hinges on predictive accuracy. Instead of generic estimates, specialized forecasting models are used for training (conditioning on 22 features), quantization, and real-time inference. The architecture optimizes a 24-hour horizon but only commits the current timestep, re-running every 30 to 60 minutes to absorb forecast errors and prevent the "end-of-world effect" where present decisions wreck future steps.