Related Experiment Videos
Identifying Key Determinants of Fertility Among Women in Bangladesh Using Machine Learning Tools: Evidence From the
Sadia Akter Sumi1, Md Toufik Umar1, Ahsanul Haque1
1Department of Statistics and Data Science University of Barishal Barishal Bangladesh.
Background And Aims:
Fertility is a fundamental demographic indicator with important implications. Identifying the characteristics associated with different fertility levels and developing reliable predictive models may provide valuable evidence for population and reproductive health planning. This study aimed to examine the factors associated with fertility levels among Bangladeshi women and to compare the predictive performance of multiple machine-learning models.
Methods:
Data were obtained from the nationally representative Bangladesh Demographic and Health Survey (BDHS) 2022. The analysis included women aged 15-49 years, with fertility categorized into three groups: no children, 1-3 children, and 4 or more children. After excluding observations with missing information on variables of interest, 9382 women were retained for the descriptive and bivariate analyses. Associations between fertility level and selected predictors were assessed using Rao-Scott design-adjusted chi-square tests. For predictive modeling, nine classification algorithms were evaluated. An 80:20 stratified train-test split and stratified five-fold cross-validation were employed, with preprocessing and SMOTE conducted within the resampling framework. Model performance was evaluated using accuracy, Cohen's kappa, Matthews correlation coefficient (MCC), area under the receiver operating characteristic curve (AUC-ROC), mean log loss, precision, recall, and F1-score. Feature importance was subsequently examined for the best-performing ensemble models.
Results:
Fertility level was significantly associated with nearly all examined characteristics except the sex of the household head. Key differences were observed by education, wealth, age at first marriage, child-death experience, reproductive preferences, and contraceptive use. LightGBM, XGBoost, and Random Forest showed the best predictive performance. LightGBM achieved the highest F1-score (0.80) and lowest log loss (0.46). Feature-importance analysis consistently highlighted husband's desire for children, child-death experience, contraceptive use, ideal number of children, husband's education, household size, and media exposure as important predictors.
Conclusion:
Fertility classification in Bangladesh is strongly associated with reproductive preferences, socioeconomic conditions, and demographic characteristics. LightGBM, XGBoost, and Random Forest demonstrated superior predictive performance, highlighting the potential of ensemble machine-learning approaches for identifying important fertility-related factors and supporting evidence-based population and reproductive health policies.
Related Concept Videos
Applications of Life Tables
Steps in Outbreak Investigation
Regression Toward the Mean