Centralized scheduling reduces regret in serial dictatorship matching problems
Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
Computer Science and Game Theory
Summary
When multiple players choose items in a fixed order, trying out one choice to learn its value can cause regret for others waiting their turn. The paper by the authors studies how to balance this learning so that the total regret across players is minimized, using a mathematical model assuming known priorities and Gaussian noise. They find a precise way to describe all possible regret outcomes and show how the scheduling of learning actions affects regret differently even under the same constraints. They also develop policies that can reach optimal trade-offs in regret without assuming a unique solution.
What this means in practice
- •For resource schedulers: Optimize scheduling in systems where resource allocation follows fixed priority orders to minimize overall learning regret among users.
- •For online marketplaces: Improve matching and exploration strategies when assigning items to users in priority order to balance learning quality and fairness.
A theory result. No direct application yet.
Authors
Lishang Xu, Guodong Ma, Pengcheng Weng, Zixuan Xia
Abstract
Exploration in centralized serial-dictatorship matching bandits must use complete matchings, so learning one player--arm pair can impose regret on others. We study this externality under a known common priority order and Gaussian rewards with unit variance. We show that the matching-level Graves--Lai constraints reduce to finitely many pairwise exploration quotas and, at top-choice-separated instances, yield a polynomial-size marginal linear program. At these instances, the exact attainable set of expected logarithmic regret coefficients is $G(θ)\Xset(θ)$, where $\Xset$ is the feasible matching-allocation set and $G$ maps allocations to player regret. The usual upper-closed Graves--Lai region can be strictly larger despite having the same Pareto-minimal boundary. We further show that identical exploration quotas can induce very different regret through their scheduling. Finally, we construct estimate--solve--track policies, uniformly good on the full row-strict class, that attain every fixed positively weighted optimum without assuming optimizer uniqueness. Every Pareto-minimal point is pointwise attainable, possibly through an instance-calibrated target.