Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

2026-08-03Information Retrieval

Information RetrievalMachine Learning
AI summary

The authors discuss how evaluating if search results truly match what users want is important but normally takes a lot of time and money because humans have to check. They created a way to use Visual Language Models (VLMs) to automatically judge relevance in Pinterest Search experiments. Their work shows that these machine-made judgments agree well with human ratings and can speed up testing significantly. This method also helps test more queries and get better quality results from experiments.

Relevance EvaluationPersonalized SearchUser Engagement MetricsHuman AnnotationVisual Language ModelsA/B TestingPinterest SearchMinimum Detectable EffectsQuery SamplingAutomated Labeling
Authors
Han Wang, Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath, Xi Chen, Roberto Konow, Kurchi Subhra Hazra
Abstract
Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that search results align with user queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present a VLM-based automated relevance evaluation pipeline deployed within Pinterest Search for online A/B experiments. We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.