A comparative study of ordinal logistic regression and machine learning models for predicting women's malnutrition in
Umme Kulsum1, Ahsanul Haque1, Pallab Barai1
1Department of Statistics and Data Science, University of Barishal, Barishal, 8254, Bangladesh.
Abstract:
Malnutrition, including both undernutrition and overnutrition, remains a major public health concern in Bangladesh, particularly among women of reproductive age. This study aims to identify key determinants of women's malnutrition in Bangladesh and compare the predictive performance of ordinal logistic regression and machine learning methods for predicting women's malnutrition using data from the 2022 Bangladesh Demographic and Health Survey. This study utilized data from 8,728 ever-married women aged 15-49 years extracted from the BDHS 2022. Six ML algorithms, including Random Forest, Extreme Gradient Boosting (XGBoost), Support Vector Machine, Naïve Bayes, AdaBoost, and Multilayer Perceptron (MLP), were compared with ordinal logistic regression by evaluating their performances using accuracy, precision, recall, [Formula: see text] score, Cohen's kappa, and area under the curve (AUC). Data preprocessing included SMOTE to address class imbalance, and models were assessed using stratified k-fold cross-validation. Findings of Ordinal Logistic Regression (OLR) suggest that age, division, residence, wealth index, current breastfeeding status, husband's education, currently working, and age at first marriage are the significant predictors of women's malnutrition. However, its predictive performance was modest, with an accuracy of 49% and macro-averaged [Formula: see text] score was 0.47. In contrast, ML models outperformed OLR across all evaluation metrics. Random Forest and XGBoost achieved the highest test accuracy (64%), with Random Forest attaining a macro-averaged [Formula: see text] score of 0.64 and achieved 66.2% accuracy (10-fold CV). Traditional models, such as OLR, are more explainable, but machine learning models demonstrate higher accuracy in classifying malnutrition. The findings can help policymakers and health professionals prioritize resources and plan targeted nutrition programs, considering the risk factors identified in this study, to lessen the burden of both undernutrition and overnutrition among women in Bangladesh.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Mechanistic Models: Compartment Models in Individual and Population Analysis
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
