Unveiling predictive factors for household-level stunting in India: A machine learning approach using NFHS-5 and
Prashant Kumar Arya1, Koyel Sur2, Tanushree Kundu3
1Institute for Human Development, Delhi, India; ICSSR Post-Doctoral Fellow, Central University of Jharkhand, Ranchi, India.
Objectives:
Childhood stunting remains a significant public health issue in India, affecting approximately 35% of children under 5. Despite extensive research, existing prediction models often fail to incorporate diverse data sources and address the complex interplay of socioeconomic, demographic, and environmental factors. This study bridges this gap by employing machine learning methods to predict stunting at the household level, using data from the National Family Health Survey combined with satellite-driven datasets.
Methods:
We used four machine learning models-random forest regression, support vector machine regression, K-nearest neighbors regression, and regularized linear regression-to examine the impact of various factors on stunting. The random forest regression model demonstrated the highest predictive accuracy and robustness.
Results:
The proportion of households below the poverty line and the dependency ratio consistently predicted stunting across all models, underscoring the importance of economic status and household structure. Moreover, the educational level of the household head and environmental variables such as average temperature and leaf area index were significant contributors. Spatial analysis revealed significant geographic clustering of high-stunting districts, notably in central and eastern India, further emphasizing the role of regional socioeconomic and environmental factors. Notably, environmental variables like average temperature and leaf area index emerged as strong predictors of stunting, highlighting how regional climate and vegetation conditions shape nutritional outcomes.
Conclusions:
These findings underline the importance of comprehensive interventions that not only address socioeconomic inequities but also consider environmental factors, such as climate and vegetation, to effectively combat childhood stunting in India.
Related Concept Videos
Nature and Nurture
Steps in Outbreak Investigation
Survival Tree
Building a Survival Tree
Constructing a...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Outliers and Influential Points


