Generative retrievers keep natural language ability while learning document ids
Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak?
Information Retrieval
Summary
Retrieval models that find documents by generating special codes often lose their natural way of talking after fine-tuning. The authors present SpeakGR, a method that trains these models to remember how to generate language while learning to retrieve documents accurately. They also introduce Adaptive SpeakGR, which adjusts the balance between learning retrieval and preserving language automatically. Their approach greatly reduces language distortion while maintaining good retrieval performance across multiple language models.
What this means in practice
- •For search engine developers: Build retrieval systems that both accurately find documents and generate fluent natural language responses without losing original language abilities.
- •For virtual assistant builders: Create assistants that can retrieve relevant information and maintain natural conversation fluency simultaneously by using dual-objective training methods.
Authors
Junchen Fu, Kleomenis Katevas, Vandana Rajan, Sofía Celi, Hamed Haddadi
Abstract
Generative retrieval (GR) enables end-to-end retrieval by generating document semantic identifiers (SIDs). However, retrieval-only fine-tuning can over-specialize pretrained language models to SID prediction, substantially distorting their natural-language distribution and limiting their suitability for interactive systems that must both retrieve documents and generate natural-language responses. We introduce SpeakGR, a dual-objective framework that learns SIDs while preserving language generation. It combines supervised SID learning with speak-preserving regularization: an on-policy distillation objective that aligns the current model with a frozen copy of the original model on student-generated prefixes using forward KL over the original text vocabulary. We further propose Adaptive SpeakGR, which dynamically adjusts the preservation strength based on observed language drift. Compared with SFT-only, SpeakGR reduces WikiText-2 forward KL by 81.3-93.8% on MS MARCO and 81.2-85.2% on Natural Questions (NQ) while retaining effective retrieval across three different LLMs. Adaptive SpeakGR further improves retrieval over SpeakGR in most settings while maintaining substantially lower language drift than SFT-only.