Learnable price randomization improves buyer surplus against adaptive optimization

Learnable Randomization as Commitment Against Adaptive Optimizers

Computer Science and Game TheoryArtificial Intelligence

Summary

This paper looks at how systems that set prices, recommend items, or classify users often face buyers or users who adapt to the system's choices. When a system always picks the single best option, it reveals too much about what the buyer will accept, allowing the buyer to respond and reduce the system's surplus. The authors propose mixing among several near-best options to keep actions unpredictable. This randomization can be learned and helps the system keep more surplus while still serving buyers acceptably. However, the benefit disappears if only one option is clearly best or if the system and user cooperate.

What this means in practice

  • For online retailers: Use learnable mixed pricing strategies to maintain higher profits against buyers who adapt to price changes.$Commercial implications: Enables dynamic pricing tools that better retain seller surplus when buyers respond strategically.
  • For recommendation system engineers: Design recommendation algorithms that mix among near-optimal items to prevent users from gaming the system and improve overall engagement.

Authors

Zihan Deng, Chuanzhi Xu, Xiaozhen Zhong, Haoyang Li, Junjie Huang

Abstract

A pricing page can walk the posted price up to the last amount a buyer still accepts, a recommender can hold back a better item for a barely acceptable promoted one, and a classifier can shift its boundary once applicants change their features. The system predicts the response and then picks the menu that serves its own objective, so the surplus above the user's cutoff is taken. Playing the single best action publishes that cutoff, while noise on actions the user would never take throws away payoff and teaches the platform that a worse menu is still acceptable. We study unpredictable near-optimal policies (UNOP), which mix uniformly on near-best actions that remain individually rational. The mixture is a commitment about the response. On a finite price grid, when the best sure-demand price strictly out-earns the randomized band, a seller who already knows the curve posts below the band, and the purchase that occurs is deterministic. Knowing that curve is not the same as predicting the next draw. The mixture can be learned and the optimizer can match its best response, while the user's payoff stays higher because the mixture changes which action is targeted. In pricing and in policy-aware recommendation this leaves more surplus than greedy play when the platform optimizes against the curve and more than one action is acceptable. The gain goes away under quality ranking, a singleton near-optimal set, a wrong utility estimate, or a short-horizon explorer. That is also where mixing should be turned off if the other side is trying to cooperate.