AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP

2026-08-31Computation and Language

Computation and LanguageArtificial IntelligenceComputers and Society
AI summary

The authors created AtlasNLP, a big collection of over 13,000 language datasets that keeps track of which countries the data represents and where it was made. They found that some countries and tasks have lots of data while others have very little. Also, just because a language is covered doesn’t mean all countries speaking it are represented. Their work highlights the need for better information about dataset geography to make language technology fairer and more accurate.

NLP datasetsgeographic metadatadata representationlanguage coveragedataset documentationdataset biasnatural language processingdataset collection
Authors
Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai, Bontu Fufa Balcha, Zayd Bashir, Angana Borah, Zara Burzo, Yubin Choi, Naihao Deng, Samika Gupta, Michel Faloughi, Claude Kwizera, Ziqiao Ma, Cynthia Yacel Fuertes Panizo, Ellie Seehorn, Hui Shen, Jiayi Tang, Zesen Zhao, Boyuan Zheng, Rada Mihalcea
Abstract
Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.