Machine learning predicts childhood stunting with attention to fairness and time

Machine Learning-Based Prediction of Childhood Stunting in Bangladesh: Fairness and Temporal Robustness Assessment

Machine Learning

Summary

Childhood stunting is a big health problem in Bangladesh caused by many factors. The authors used data from 2007 to 2022 to train computer programs that predict which children might be stunted. They tested how well these programs worked over different years and for different groups of children, like boys versus girls or city versus rural kids. Their study shows that prediction accuracy changes over time and among groups, highlighting the need to check fairness and reliability when using such tools in health decisions.

What this means in practice

  • For public health planners: Use machine learning models validated over time and demographic groups to inform programs targeting childhood stunting in Bangladesh.
  • For data scientists in health: Develop stunting risk prediction tools incorporating model fairness checks and temporal validation for use in similar health monitoring datasets.

Authors

Md Ahshanul Haque, Muhammad Ashad Kabir

Abstract

Childhood stunting remains a major public health concern in Bangladesh and reflects long-term growth failure influenced by child, maternal, household, socioeconomic, and health-service factors. This study used nationally representative Bangladesh Demographic and Health Survey data from 2007 to 2022 to develop machine learning models for population-level prediction of childhood stunting and to assess temporal robustness and subgroup fairness. Children aged 0-59 months with complete anthropometric and predictor data were included. Data from the 2007, 2011, and 2014 survey rounds were used for model development, while the 2018 and 2022 rounds were retained as temporal test datasets. Twelve feature-selection approaches were assessed, and the KNN permutation importance-selected predictor set was used for final model evaluation. Eleven machine learning models were evaluated: ten conventional algorithms and one pretrained tabular foundation model, TabPFN. Performance was assessed using balanced accuracy, AUROC, F1-score, Brier score, and expected calibration error. Subgroup fairness was examined by child sex, place of residence, and socioeconomic status. The final analytic sample included 18,844 children, of whom 35.05% were stunted. In the development hold-out test dataset, TabPFN showed the highest observed balanced accuracy overall at 67.58%, while AdaBoost showed the highest observed balanced accuracy among conventional models at 67.51%. In temporal testing, the highest observed balanced accuracy was found for Gradient Boosting in BDHS 2018 and XGBoost in BDHS 2022. Model performance varied across survey rounds and subgroups, highlighting the importance of temporal validation, subgroup fairness assessment, and transparent interpretation in public health prediction modeling.