New method learns optimal choices in two-arm bandit problems faster
Elicitation and Decision Geometry in Single-Index Bandits
Machine Learning
Summary
This paper deals with a decision-making problem where you have two options, each with unknown rewards that depend on context. The authors develop a new method called Natural Boundary Learning (NBL) that learns the best choice directly without needing to estimate how each option behaves separately. They show that this approach quickly improves decisions over time by focusing on the boundary between the two options. Their analysis explains when and why the method works well, and experiments confirm these findings.
What this means in practice
- •For online recommendation engineers: Improve personalization by efficiently learning when to switch between two content options based on user context without modeling each option fully.
- •For adaptive clinical trial designers: Optimize patient treatment assignment between two therapies by learning decision boundaries directly from patient features, reducing need for complex reward modeling.
Authors
Sakshi Arya, Cheng Soon Ong
Abstract
We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (NBL), a greedy procedure that uses a sequential Stein contrast to learn the optimal boundary directly, without estimating the reward functions or the common link. We characterize the local Riemannian dynamics of NBL through a decision stability coefficient balancing arm separation, link geometry, and the context distribution. We show that this stability is connected to the elicitation geometry of the underlying convex potential. Under local decision stability, NBL contracts toward the optimal boundary and achieves $O(\log n)$ expected regret. Numerical experiments illustrate the predicted stability regimes and compare NBL with a parametric greedy benchmark under link misspecification.