New click models improve ranking evaluation for recommendation systems

Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition

Information RetrievalMachine Learning

Summary

Recommendation systems need to estimate how well new ranking strategies will perform using past user click data. However, clicks alone can’t tell if an item was shown but ignored or never shown at all, leading to errors. The authors propose new methods that separate the chance an item was seen from its actual relevance to the user. Their improved techniques reduce bias in evaluating recommendations, especially when item visibility depends on its position in the list. These methods work best with large data sets and have some limits if user behavior is very sequential.

What this means in practice

Authors

Riki Okamura, Toshiharu Sugawara

Abstract

Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We propose two estimators based on the decomposition of clicks into examination and relevance. First, the latent-examination independent inverse propensity score (LE-IIPS) estimator corrects the IIPS bias using policy examination probability ratios. Second, the examination-decomposed doubly robust (ED-DR) estimator extends LE-IIPS to a doubly robust framework. ED-DR is unbiased if the examination probabilities are correct regardless of relevance accuracy, or under ranking-independent examination, even if both model estimates are inaccurate. Experiments show that ED-DR achieves a lower MSE than existing methods with large sample sizes, especially when the examination depends on ranking. We also highlight its limitations under small samples or cascade user behavior conditions.