使用机器学习技术来预测巴西统计学课程的本科生学率
Renata Rojas Guerra1, Marcos Antonio DE Azevedo DE Campos1, Andressa Lopez Soares1
1Federal University of Santa Maria, Department of Statistics, Roraima Avenue, 1000, 97105-900 Santa Maria, RS, Brazil.
Anais da Academia Brasileira de Ciencias
|December 19, 2025
概括
在巴西统计学课程中,学生学率很高 (65%). 后勤回归准确地预测结果,将入学时间确定为关键风险因素,而社会支持和活动可以降低退学率.
科学领域:
- 教育数据挖掘教育数据挖掘
- 高等教育中的机器学习
- 统计学 教育 研究 研究 研究
背景情况:
- 巴西统计学本科课程的高学生学率是一个重大挑战.
- 了解影响学生退学的因素对于开发有效干预措施至关重要.
研究的目的:
- 开发和评估机器学习模型,用于在巴西统计计划中对学生学进行分类.
- 在这个学术背景下,确定与学生退学相关的关键预测因素.
主要方法:
- 利用了来自国家教育研究与研究研究所 (ANÍSIO TEIXEIRA (INEP)) 高等教育普查的微数据.
- 分析了10387名2009年至2014年的本科生数据,这些学生被监测到2017年.
- 应用并比较了物流回归,随机森林和支持矢量机算法来进行退学分类.
主要成果:
- 后勤回归模型实现了最高的准确度 (86.1%),并提供了更好的解释性.
- 招生时间被确定为学生学的主要预测因素.
- 社会支持和参与互补活动显著降低了学的可能性.
结论:
- 机器学习,特别是逻辑回归,为预测学生在统计学课程中学提供了一种有效的方法.
- 专注于学生参与,社会支持和课程持续时间的干预措施可以降低退学率.
相关概念视频
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Random Sampling Method
14.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
14.0K
Outliers and Influential Points
5.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.9K
Regression Analysis
7.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.8K
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K
Censoring Survival Data
502
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
502
