The Illusion of Cross-Lingual Safety in Low-Resource Languages
2026-08-11 • Computation and Language
Computation and Language
AI summaryⓘ
The authors studied how well safety features in large language models (LLMs) work across different languages, focusing on four African languages. They found that safety signals learned in English do not transfer well to these languages, meaning harmful prompts often aren't blocked effectively. Even when prompts are translated or culturally adapted, the models recognize their meaning but fail to enforce safety. This suggests that current multilingual safety efforts might only work superficially and not truly protect users in low-resource languages.
large language modelssafety alignmentcross-lingual transferlow-resource languagesAfrican languageslatent representationsmodel hidden statesprompt localizationmultilingual NLPharmful content detection
Authors
Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay, Maryam Ibrahim Mukhtar, Esmael Ahmed Abdu, Tassallah Abdullahi, Jessica Oparebea, Saminu Mohammad Aliyu, Idris Abdulmumin, Abubakar Juma Chilala, Nicholaus Dismas Ladislaus, Alfred Malengo Kondoro, Lemofouet Valdini Douglace, Shamsuddeen Hassan Muhammad, Seid Muhie Yimam
Abstract
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.