Method improves diversity of AI text generation without trade offs
Improving the Diversity of LLM Outputs without a Trade-off
Computation and LanguageArtificial IntelligenceMachine Learning
Summary
Language models sometimes produce similar or repetitive responses, limiting the variety of ideas they generate. The authors propose a method called DAST that rearranges how tokens (words or pieces of words) are ordered internally so that similar meanings appear together. This change helps make similar tokens less likely to be picked repeatedly, increasing output diversity without slowing down generation or changing what the model thinks is likely overall. Their approach improves results on a question-answering task, showing it can produce a wider range of good answers.
What this means in practice
- •For natural language developers: Generate more diverse and creative text outputs from language models without increased computational cost or reduced reliability.
- •For conversational ai teams: Improve response variety in chatbots and virtual assistants to enhance user engagement by using token reordering with arithmetic sampling.
Authors
Ryoma Sato
Abstract
We propose DAST (Diversifying Arithmetic Sampling with TokenTour), a method that increases the diversity of LLM outputs without any change to the marginal distribution and with negligible generation-time overhead (a few microseconds). We observe that token IDs are often arranged in a meaningless order and reassign them so that tokens with similar meanings appear consecutively. This can be done in advance in a few hundred seconds per model, and the resulting order can be reused for all subsequent generations. By combining this order with arithmetic sampling (or quasi-Monte Carlo methods), we make similar tokens less likely to be generated across runs while preserving the distribution. Our method not only produces qualitatively good ideas but also significantly improves performance on the downstream task of ProtoQA.