Inference auction lets users bid for faster model responses

Inference Auctions

Machine LearningArtificial IntelligenceComputer Science and Game Theory

Summary

When lots of people ask a big AI model for answers at the same time, it can’t respond to everyone quickly. The authors created a way for people to bid money to get their answers faster, so the fastest answers go to those who value speed most. They built smart software that sets prices fairly and helps users bid within their budgets. Tests show this bidding system makes the whole AI-serving process work better without slowing it down.

What this means in practice

  • For cloud ai platform operators: Set prices dynamically for AI model usage to allocate limited compute resources efficiently and improve overall system responsiveness.$Commercial implications: Enables selling faster AI inference access with optimized pricing that increases revenue from priority service.
  • For real-time ai service teams: Manage competing user requests by prioritizing workloads based on user bids to maintain low latency even under heavy load.

Authors

Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan

Abstract

When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing latency, and we develop fast algorithms for implementing prices that incentivize truthful bidding. We also design an autobidding agent for our inference auction, where users specify an inference budget and the autobidder dynamically adjusts its bids over time to maximize user utility subject to the budget constraint. Experiments validate the practicality of our auction: it increases system welfare while maintaining the cache utilization and latency advantages of SGLang, a state-of-the-art inference serving framework.