RenderRank speeds up document reranking using visual token compression
RenderRank: Learning to Rerank Text with Compressed Visual Tokens
Information Retrieval
Summary
When computers search through many documents to find the best ones for a question, they usually look at the text directly. This can take a lot of time because each word is processed. The authors show that by turning documents into small pictures and using a computer vision approach, they can look at fewer pieces and still decide which documents are the best matches. This way, the process is faster and sometimes even more accurate than traditional text-based methods.
What this means in practice
- •For search engine developers: Use compressed visual token representations to speed up multi-document relevance reranking while maintaining or improving accuracy.
- •For digital library maintainers: Handle long documents more efficiently by reranking with visual tokens that reduce token counts and processing time.
Authors
Seongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon, Heuiseok Lim
Abstract
Rendering document text as images allows vision-language models to encode documents as visual tokens, which can reduce input sequence length compared with text input. This reduction in input length is particularly useful for reranking, where each query involves scoring multiple candidate documents and token savings apply to each candidate evaluation. We introduce RenderRank, a reranker that learns query-dependent relevance scoring from compressed visual document representations instead of the text token sequences used by conventional text-based rerankers. Training first aligns relevance scores from visual inputs with those of a text-based teacher, then refines the relative scores of positive and negative documents for the same query. Across 11 datasets from BEIR, RenderRank uses 16.5-35.5% fewer input tokens while achieving an average NDCG@10 of 55.96, outperforming all evaluated text-based baselines below 4B parameters and some larger models. Across four long-document datasets, it achieves an average NDCG@10 of 88.27 with approximately half the average input token count of the evaluated text-based rerankers. In this setting, RenderRank delivers 1.70x the highest average throughput of the evaluated baselines. These results demonstrate that compressed visual representations can support accurate document relevance scoring, providing an alternative to text token representations for reranking.