Papers for

multilingual product managers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Large language models show uneven cultural preferences and resist prompt corrections

DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs

Abstract: Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single "correct" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.

Wed 9 SeptComputation and LanguageArtificial Intelligence
The gist
Large language models (LLMs) that power assistants often favor certain cultures, especially US and UK, when asked to respond to everyday situations. This bias can affect how users from different cultures experience these systems. The authors created DiSCo, a new way to measure these cultural preferences and how much they can be changed by prompts or added facts. They found that usual prompts make the bias worse and adding cultural facts does little to change the models' preferences. This shows that fixing cultural bias in language AIs is more complicated than just guiding them with prompts.
Open 2609.10253v1