Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
2026-08-12 • Computation and Language
Computation and LanguageArtificial IntelligenceComputers and Society
AI summaryⓘ
The authors explain that AI tools meant to help education in under-resourced areas often overlook languages like Bengali, even though many people speak them. They point out four main problems: very little Bengali content online, less training data compared to English, challenges with how Bengali script is broken into parts for AI processing, and less internet access in rural Bengali areas. These issues come from bigger decisions about where resources go and how AI systems are built. The authors suggest that the shortage of data is a bigger, systemic problem, and that creating AI tools that work offline could help make things fairer.
Artificial IntelligenceBengali languageTraining corporaTokenizationMultilingual modelsData scarcityOffline-first designInternet penetrationResource allocationEquity in AI
Authors
Avijit Roy, Proma Roy
Abstract
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.