Related Experiment Video
Updated: Jan 8, 2026

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
Using machine learning techniques for predicting the dropout of undergraduate students in Brazilian courses of
Renata Rojas Guerra1, Marcos Antonio DE Azevedo DE Campos1, Andressa Lopez Soares1
1Federal University of Santa Maria, Department of Statistics, Roraima Avenue, 1000, 97105-900 Santa Maria, RS, Brazil.
Abstract:
This research aims to propose a machine learning approach to classify dropout outcomes among students in Statistics undergraduate programs in Brazil, identifying the most important factors associated with this phenomenon. This study uses microdata from the National Institute for Educational Studies and Research Anísio Teixeira (INEP) Higher Education Census to analyze the experiences of 10,387 undergraduate students who began their bachelor's degrees between 2009 and 2014 in Brazil and were monitored until the end of 2017. Approximately 65% of these students dropped out, highlighting the critical dropout challenge in Brazilian Statistics programs. The outcome of interest is the student's final status in the census data, either completion or dropout. Logistic regression, random forest , and support vector machine algorithms were fitted in the learning stage. For each algorithm, the best model was evaluated on test data. Results indicate that the logistic regression model outperforms the others, achieving an accuracy of 86.1% and offering greater interpretability compared to the other models. The study reveals that the duration of enrollment is a key predictor of dropout likelihood. Moreover, receiving social support and participating in complementary activities are significant factors in reducing dropout rates.
Related Concept Videos
Regression Toward the Mean
Random Sampling Method
Outliers and Influential Points
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Censoring Survival Data