Multilingual medical AI faces debate on consistency versus cultural adaptation

Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions

Computation and Language

Summary

When large language models answer medical questions in different languages, there is a debate about whether their answers should always be the same or change to fit cultural differences. The authors looked at studies and surveyed experts in medicine, natural language processing, and anthropology in several countries. They found that anthropologists usually support adapting answers to culture, while medical and AI experts are split, especially between the US and Europe. The AI models tested did not show this cultural sensitivity and tended to give consistent answers. The study shows that more research is needed to understand which approach works better for people in different cultures.

large language modelsmultilingual NLPmedical question answeringcross-lingual consistencycultural adaptationbenchmarkingstakeholder perspectivesnatural language processingmedical professionalsanthropology

Authors

Minh Duc Bui, Mario Sanz-Guerrero, Abteen Ebrahimi, Sagi Shaier, Peter Herbert Kann, Manuel Mager, Katharina von der Wense

Abstract

Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.