Thompson sampling method improves bandit decisions with changing baselines
Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits
Machine Learning
Summary
This paper looks at how computers decide between multiple options when the overall success rate changes over time. The usual methods remember each option’s absolute performance, which can get outdated if conditions shift. The authors propose a new method called Odds-Ratio Thompson Sampling that focuses on comparing differences between options instead, refreshing the shared baseline regularly. Their tests show this new method makes better decisions when conditions vary and avoids misleading memory problems. However, when the differences between options themselves change a lot, this method does not help.
What this means in practice
- •For digital marketers: Run A/B tests that remain reliable even when overall user engagement changes over time by using a contrast-based update method.
- •For online platform engineers: Improve adaptive content allocation algorithms by tracking relative differences between options rather than absolute success metrics.
Authors
Sulgi Kim
Abstract
Batched multi-armed bandits update on a service's own schedule, and the usual implementation carries each arm's absolute reward rate from one update to the next. When the shared level moves between batches, that memory goes stale even though the comparisons between arms may not have. Odds-Ratio Thompson Sampling (OR-TS) instead carries the joint posterior over log-odds contrasts and fits the common level afresh in every batch, marginalizing it out. This paper specifies that update, places it inside a Bayesian bandit agent with two controls, decay for how much past evidence survives an update and aggressiveness for how sharply belief becomes allocation, and evaluates it against absolute-rate memory. Across 86 public A/B series the level varies about twenty-five times more than the contrast. In prespecified synthetic environments a moving level costs absolute-rate memory five times the regret and leaves the best arm below a majority of traffic in 7 of 20 runs, against none for OR-TS. In a policy simulation built from 71 real experiments, where the contrasts are too small to resolve, expected-click differences stay within 0.1% for 58 of them, yet contrast memory still ends on the better arm more than twice as often. Where the contrasts themselves move, the bet fails, and that case is reported too.