SPADE metric measures truly surprising recommendations beyond popularity and similarity

SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations

Information RetrievalArtificial IntelligenceMachine Learning

Summary

Recommender systems often suggest items based on popularity or similarity, which can make recommendations predictable. The authors created SPADE, a new way to evaluate how surprising or serendipitous recommendations are by considering popularity, similarity, and what the user actually likes. SPADE works by mapping items on a chart and measuring how far recommended items are from popular or similar items that a user already knows. Testing this across several datasets showed SPADE reliably identifies recommendations that are both relevant and genuinely unexpected.

What this means in practice

  • For online retail teams: Evaluate product recommender systems to balance popularity, similarity, and true user interest for more engaging shopping experiences.
  • For streaming service engineers: Measure how well music or video recommendation algorithms present surprising yet relevant content to users beyond standard popularity trends.

Authors

Tobias Vente, Maarten Peirsman, Noah Daniëls, Hannu Toivonen, Bart Goethals

Abstract

Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.