Sign aware recommender systems need better evaluation metrics

What Gets Measured Gets Managed: Sign-aware Recommendation Needs Sign-aware Evaluation

Information Retrieval

Summary

Many recommendation systems try to guess what you like, but some also use what you dislike to improve suggestions. The authors found that current methods that should use dislikes still often show things you don’t like in top recommendations because their scoring doesn’t properly handle negative feedback. This problem is hidden because usual tests don’t penalize recommending disliked items. The authors propose new ways to measure how well systems avoid disliked items, helping improve recommendation quality.

What this means in practice

  • For recommender system engineers: Improve recommendation quality by using sign-aware evaluation metrics to reduce disliked content in user top-K results.
  • For online retail teams: Reduce customer dissatisfaction by integrating sign-aware metrics that explicitly penalize disliked item recommendations in product suggestions.

Authors

Minchan Kim, Jungmin Hwang, Hyunwoo Park

Abstract

Sign-aware recommender systems have recently been developed to leverage negative feedback for a deeper understanding of user preferences. However, our empirical diagnosis reveals that state-of-the-art graph-based sign-aware recommender systems are paradoxically valence-blind. Even though they explicitly incorporate sign information during training, they consistently fail to differentiate liked items from disliked ones at the ranking stage, frequently infiltrating top-K recommendations with disliked content. Through linear probing, we show that while valence information exists in the learned embeddings, it remains inaccessible to the inner-product scoring function. This widespread failure remains entirely undetected because conventional evaluation metrics, such as Recall, HR, and NDCG, assign a uniform utility of zero to both negative and unobserved items, creating a systematic evaluation blind spot. To bridge this gap, we propose a family of signed metrics, Signed Recall, Signed HR, and Signed NDCG, that explicitly penalize the recommendation of disliked content. Systematic re-evaluation under our proposed metrics fundamentally reshapes the established performance landscape, revealing that methods ranked highly under conventional metrics often fail to protect users from disliked content. Finally, through a proof-of-concept auxiliary loss, we confirm that the proposed metrics provide actionable training signals, guiding models toward valence-aware behavior without sacrificing conventional relevance. For transparency, our source code is available at: https://anonymous.4open.science/r/signed-rec-benchmark-07E4