利用机器学习,分析影响大学生学业成绩的因素
Yuri Reina Marín1, Lenin Quiñones Huatangari2, Judith Nathaly Alva Tuesta1
1Oficina de Gestión de la Calidad, Universidad Nacional Toribio Rodríguez de Mendoza de Amazonas, Chachapoyas , Peru.
Scientific reports
|November 27, 2025
概括
机器学习通过分析各种因素,准确地预测大学生学术表现. 整合多个学生维度可以提高模型的概括性,从而获得更好的教育见解.
科学领域:
- 教育数据挖掘教育数据挖掘
- 机器学习在教育中的应用
- 高等教育研究 高等教育研究
背景情况:
- 评估大学生的学业成绩是复杂的,因为不同的影响因素和非线性学习模式.
- 高等教育机构在各种社会,经济和学术环境中优先了解学术表现的决定因素.
- 传统的评估方法在捕捉学生成功的多面性质方面面临挑战.
研究的目的:
- 开发和评估用于预测大学生学术表现的机器学习模型.
- 确定影响学术成功的关键决定因素,跨人口,社会经济,学术和心理社会层面.
- 评估各种教育数据挖掘算法的有效性,以建模复杂的学生行为.
主要方法:
- 利用了386名大学生的主要数据.
- 评估了九种机器学习算法:XGBoost,随机森林,ANN,SVM,决策树,naive Bayes,物流回归,AdaBoost和KNN.
- 分析因素包括人口统计,社会经济地位,学术界,社会/家庭生活,健康,基础设施和时间管理.
主要成果:
- 机器学习有效地模拟了学业成绩与其决定因素之间的非线性关系.
- 模型准确地预测了学术成果,敏感性和可解释性分析解释了个别变量贡献.
- 集成多个学生维度的算法显示出优越的概括能力.
- 所有评估的算法在建模教育动态方面都表现出有效性,尽管预测准确度有所不同.
结论:
- 机器学习为预测和理解大学生学术表现提供了强大的工具.
- 在不同的大学环境中进行准确的评估需要将多种影响因素纳入预测模型.
- 未来的研究应该利用全面的数据和先进的机器学习技术来进行强大的教育分析.
相关概念视频
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Correlations
35.7K
Correlation means that there is a relationship between two or more variables (such as ice cream consumption and crime), but this relationship does not necessarily imply cause and effect. When two variables are correlated, it simply means that as one variable changes, so does the other. We can measure correlation by calculating a statistic known as a correlation coefficient. A correlation coefficient is a number from -1 to +1 that indicates the strength and direction of the relationship between...
35.7K
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Theory of Attribution II: Kelley's Covariation Theory
465
Attribution theory plays a crucial role in social psychology, helping to explain how individuals interpret the causes of behavior. One prominent model within this field is Harold Kelley's covariation theory, which provides a systematic approach to determining whether internal traits or external circumstances drive a person's actions. The model posits that individuals rely on three key types of information—consensus, consistency, and distinctiveness—to make these judgments.Consensus:...
465
Mechanistic Models: Compartment Models in Individual and Population Analysis
226
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
226

