Multilingual data mixing improves reasoning in language models
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
Computation and Language
Summary
Most language models are better at reasoning in English than other languages, which can be a problem for people who speak different languages. The authors worked on making a model reason well in the same language it is asked in, across 60 languages. They found that mixing a wide variety of languages and types of data helps models learn to think in many languages without needing reasoning examples in every one. This means models can understand and answer questions more naturally in different languages. The authors also shared their model and data to support further development.
What this means in practice
- •For multilingual chatbot developers: Create chatbots that reason accurately in the user's language across numerous languages without needing separate reasoning training data for each one.
- •For multilingual content generation teams: Improve automated content generation systems to produce culturally sensitive and language-consistent responses by training on mixed multilingual data.
Authors
Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza, Alexandre Berard, Thomas Euyang, Marzieh Fadaee, Julia Kreutzer
Abstract
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in. This is inaccessible for non-English-speaking users, risks losing the intent of the original question, and forgoes knowledge more readily expressed in the target language. In this work, we advance L2 reasoning, the ability of a model to reason consistently in the language of the user's prompt, thus building an in-language bridge between the prompt and the answer. We approach this problem from a data-centric angle, investigating how to optimize data composition and scheduling in SFT for reasoning generalization. Building Tiny Aya L2-Thinker at 3.35B scale, we achieve an L2 reasoning rate above 93% across 60 languages on 6 benchmarks spanning math, commonsense reasoning, instruction following, open-ended generation, and cultural reasoning while keeping performance strong. We show the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone. These findings indicate that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing and without requiring reasoning supervision in every target language. We release our model weights and multilingual reasoning data to support further research on accessible, in-language reasoning.