Papers for

online retail teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Sign aware recommender systems need better evaluation metrics

What Gets Measured Gets Managed: Sign-aware Recommendation Needs Sign-aware Evaluation

Abstract: Sign-aware recommender systems have recently been developed to leverage negative feedback for a deeper understanding of user preferences. However, our empirical diagnosis reveals that state-of-the-art graph-based sign-aware recommender systems are paradoxically valence-blind. Even though they explicitly incorporate sign information during training, they consistently fail to differentiate liked items from disliked ones at the ranking stage, frequently infiltrating top-K recommendations with disliked content. Through linear probing, we show that while valence information exists in the learned embeddings, it remains inaccessible to the inner-product scoring function. This widespread failure remains entirely undetected because conventional evaluation metrics, such as Recall, HR, and NDCG, assign a uniform utility of zero to both negative and unobserved items, creating a systematic evaluation blind spot. To bridge this gap, we propose a family of signed metrics, Signed Recall, Signed HR, and Signed NDCG, that explicitly penalize the recommendation of disliked content. Systematic re-evaluation under our proposed metrics fundamentally reshapes the established performance landscape, revealing that methods ranked highly under conventional metrics often fail to protect users from disliked content. Finally, through a proof-of-concept auxiliary loss, we confirm that the proposed metrics provide actionable training signals, guiding models toward valence-aware behavior without sacrificing conventional relevance. For transparency, our source code is available at: https://anonymous.4open.science/r/signed-rec-benchmark-07E4

Sun 27 SeptInformation Retrieval
The gist
Many recommendation systems try to guess what you like, but some also use what you dislike to improve suggestions. The authors found that current methods that should use dislikes still often show things you don’t like in top recommendations because their scoring doesn’t properly handle negative feedback. This problem is hidden because usual tests don’t penalize recommending disliked items. The authors propose new ways to measure how well systems avoid disliked items, helping improve recommendation quality.
Open → 2609.33346v1

SPADE metric measures truly surprising recommendations beyond popularity and similarity

SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations

Abstract: Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.

Fri 25 SeptInformation RetrievalArtificial IntelligenceMachine Learning
The gist
Recommender systems often suggest items based on popularity or similarity, which can make recommendations predictable. The authors created SPADE, a new way to evaluate how surprising or serendipitous recommendations are by considering popularity, similarity, and what the user actually likes. SPADE works by mapping items on a chart and measuring how far recommended items are from popular or similar items that a user already knows. Testing this across several datasets showed SPADE reliably identifies recommendations that are both relevant and genuinely unexpected.
Open → 2609.31164v1

Multimodal simulation improves website user experience recommendations

Automatic multimodal UX improvement recommendations from LLM agent user simulations

Abstract: Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.

Sat 19 SeptComputation and LanguageArtificial IntelligenceHuman-Computer Interaction
The gist
Testing how easy and pleasant websites are to use is usually expensive and slow because it needs real people. The authors created a system called AMUSER that uses AI to simulate users interacting with websites, including both what they see and do. This lets the system suggest ways to improve websites automatically and faster, with better suggestions than previous approaches that looked only at text actions. Interestingly, giving the AI visual information during the testing part helps a lot, but giving visuals when making suggestions can slightly hurt quality.
Open → 2609.22971v1