Related Experiment Video
Updated: Jan 8, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Comparison of machine learning classification and regression models for prediction of academic performance among
Amira Fathy Abdallah Sayed1,2, Mostafa Ahmed Arafa1, Nessrin Ahmed El-Nimr1
1Department of Epidemiology, High Institute of Public Health Alexandria University, Alexandria, Egypt.
Abstract:
Machine learning (ML) is an artificial intelligence tool that focuses on learning by generating models using established algorithms that represent a given dataset. It can be used as a predictive tool for students' academic performance (AP) at both undergraduate and postgraduate levels. A cross-sectional analysis was conducted using academic records of 922 postgraduate students admitted to the High Institute of Public Health, Alexandria University, Egypt, between 2020-2024. Data included 22 features spanning pre-enrollment metrics, academic performance, and demographic traits. Classification algorithms, and regression models were trained on 75% of the dataset, validated via 5-fold cross-validation. Performance metrics included accuracy, precision, recall, AUC for classification, and MAE, RMSE, and R² for regression. Regression models outperformed classification models in AP prediction, with Ensemble (Soft Voting) achieving the highest accuracy (74.25%), lowest MAE (0.3383), and RMSE (0.4316). Among classification models, Random Forest demonstrated superior accuracy (71.43%) and AUC (0.87). Numerical features like the number of failed courses showed the strongest negative correlation with AP (r = -0.37). Key predictors included bachelor's university, major, department, and pre-enrollment CGPA. Feature importance analysis highlighted failed courses as the top determinant, followed by institutional and academic background variables. Regression-based ML models, particularly Ensemble (Soft Voting), proved more effective than classification approaches for predicting nuanced variations in AP. These findings enable institutions to prioritize early interventions for at-risk students, and optimize resource allocation. However, moderate R² values (0.3832) underscore the need to integrate psychosocial and behavioral factors in future studies.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Related Concept Videos
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Mechanistic Models: Compartment Models in Individual and Population Analysis
Regression Toward the Mean
Comparing the Survival Analysis of Two or More Groups