Dharma AI describes a constraint-aware GPU allocation system that improved cluster utilization by up to 33 percentage points compared to FIFO scheduling on identical hardware and workloads. The allocator treats real-time inference demand as a curve rather than a fixed reservation and sequences batch jobs by priority across a 24-hour horizon. Across seven benchmark scenarios, the system achieved average priority-weighted output gains of 52%, ranging from 15.9% to 105% depending on the workload mix.
