Papers for

virtual assistant builders

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Generative retrievers keep natural language ability while learning document ids

Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak?

Abstract: Generative retrieval (GR) enables end-to-end retrieval by generating document semantic identifiers (SIDs). However, retrieval-only fine-tuning can over-specialize pretrained language models to SID prediction, substantially distorting their natural-language distribution and limiting their suitability for interactive systems that must both retrieve documents and generate natural-language responses. We introduce SpeakGR, a dual-objective framework that learns SIDs while preserving language generation. It combines supervised SID learning with speak-preserving regularization: an on-policy distillation objective that aligns the current model with a frozen copy of the original model on student-generated prefixes using forward KL over the original text vocabulary. We further propose Adaptive SpeakGR, which dynamically adjusts the preservation strength based on observed language drift. Compared with SFT-only, SpeakGR reduces WikiText-2 forward KL by 81.3-93.8% on MS MARCO and 81.2-85.2% on Natural Questions (NQ) while retaining effective retrieval across three different LLMs. Adaptive SpeakGR further improves retrieval over SpeakGR in most settings while maintaining substantially lower language drift than SFT-only.

Mon 28 SeptInformation Retrieval
The gist
Retrieval models that find documents by generating special codes often lose their natural way of talking after fine-tuning. The authors present SpeakGR, a method that trains these models to remember how to generate language while learning to retrieve documents accurately. They also introduce Adaptive SpeakGR, which adjusts the balance between learning retrieval and preserving language automatically. Their approach greatly reduces language distortion while maintaining good retrieval performance across multiple language models.
Open → 2609.35430v1