通过可解释的建模来预测高等教育中的学术成绩
1School of Foreign Languages, Wuhan Business University, Wuhan, Hubei, People's Republic of China.
PloS one
|September 5, 2024
概括
预测学生的学业成绩对于教育质量至关重要. 新的XGB-SHAP模型准确地预测学生的成绩,识别了诸如自主学习和教学模式等关键因素.
科学领域:
- 教育技术的教育技术
- 机器学习在教育中的应用
- 高等教育中的数据科学
背景情况:
- 学生的学术成绩是教育质量的关键指标.
- 预测成绩有助于教育工作者量身定制教学和提高学生成绩.
- 使用传统方法从教育数据中提取可操作的见解存在挑战.
研究的目的:
- 引入一种新的机器学习方法,即XGB-SHAP模型,用于预测学生的学业成绩.
- 解决传统算法在界定影响学生成绩的因素方面的局限性.
- 提高高等教育中学生绩效预测的准确性.
主要方法:
- 这项研究使用了极端梯度增强 (XGBoost) 算法与夏普利添加式扩展 (SHAP) 结合.
- 该XGB-SHAP模型应用于87名大学生在日语课程中的数据集 (2021年9月 - 2023年6月).
- 模型性能与其他三种机器学习模型进行了比较.
主要成果:
- XGB-SHAP模型显示出高精度,平均绝对误差 (MAE) 大约为6和R平方值为0.82.
- 该模型在预测准确性方面表现优于其他三种机器学习模型.
- 分析揭示了不同教学模式如何影响影响学生成绩的因素.
结论:
- XGB-SHAP模型为预测学生学业成绩提供了一种卓越的方法.
- 根据教学模式进行定制的特征选择是有效预测的必要条件.
- 将自主学习技能整合到预测指标中,对于准确的绩效预测至关重要.
更多相关视频
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.1K
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
7.5K
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Correlations
32.7K
Correlation means that there is a relationship between two or more variables (such as ice cream consumption and crime), but this relationship does not necessarily imply cause and effect. When two variables are correlated, it simply means that as one variable changes, so does the other. We can measure correlation by calculating a statistic known as a correlation coefficient. A correlation coefficient is a number from -1 to +1 that indicates the strength and direction of the relationship between...
32.7K
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
