Related Experiment Video
Updated: Jan 14, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Monte Carlo optimization for sampling selection in imbalanced data applied to student dropout prediction
Dianela Herrera1, Nicolás Ángel1, Diego González1
1Departamento de Física, Universidad Católica del Norte, Av. Angamos 0610, Antofagasta 1240000, Chile.
Abstract:
University student completion rates vary among students who initially enroll in academic degrees and professional careers. Student dropout is a widespread global phenomenon that transcends both the type of degree pursued and the university attended. Traditionally, the greatest emphasis in corrective measures has been placed on improving academic performance and, to a lesser extent, on other variables that are overshadowed by the former but are equally impactful. Therefore, the current motivation is to develop an effective machine learning-based tool to identify students at a higher risk of dropping out early after 1-3 years of study. We use a large dataset from the Universidad Católica del Norte to test the methodology. Machine learning specific tools are tested to verify their predictive capability, and their results are discussed to remark on their precise utility. Moreover, we address the class imbalance in the first-year data by implementing an innovative adjustment using the Monte Carlo methodology, improving model performance under imbalanced conditions. Indeed, the technique is mainly relevant to first-year dropout, where the dataset is more anomalous. Nevertheless, a level of improvement is observed in all cases studied. The ultimate goal is to identify at-risk students early to support the timely, effective, and proper implementation of preventive interventions.
Related Concept Videos
Random Sampling Method
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Choosing Between z and t Distribution
Convenience Sampling Method
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Quantifying and Rejecting Outliers: The Grubbs Test
Systematic Sampling Method
Systematic sampling is one of the simplest methods...

