一种基于机器学习和沙普利增量解释的模型,用于预测落后地区在标准化测试中的学业成绩
Mario Suaza-Medina1,2, Rita Peñabaena-Niebles3, Maria Jubiz-Diaz2
1Department of Informatics and Computer Science, Universidad de Zaragoza, Maria de Luna 1, Zaragoza, 50018, Spain.
教育数据挖掘 (EDM) 和Shapley值揭示了影响哥伦比亚Saber 11考试学生成绩的关键因素,特别是在服务不足的地区. 社会经济地位,性别和地区显著影响学术成功.
科学领域:
- 教育教育教育教育教育教育.
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 数据分析对于提高教育质量和预测学生行为至关重要.
- 学业成绩受到地区因素的影响,如人口统计和社会经济地位,特别是在落后的地区.
研究的目的:
- 通过教育数据挖掘 (EDM) 和Shapley值来识别影响哥伦比亚Saber 11考试学生成绩的关键变量.
- 将这些方法应用于落后地区,以了解和预测学业成绩.
- 分析个别变量对学生成绩的影响.
主要方法:
- 9个分类算法的应用.
- 使用Shapley值进行变量重要性分析.
- 专注于哥伦比亚Saber 11标准化测试数据.
主要成果:
- 极端梯度提升机,轻梯度提升机和梯度提升机展示了最高的精度.
- 莎普利价值观确定的最有影响力的变量包括社会经济水平,性别,地区,机构位置和年龄.
- 来自城市机构的18岁以上的男性学生取得了最佳表现;Nariño学生显示优势.
结论:
- 该方法有效地识别落后地区的业绩预测指标.
- 调查结果可以为教育公平的公共政策提供信息.
- 在落后地区之间,教育质量存在显著差异.
更多相关视频
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
相关概念视频
Reliability and Validity
Outliers and Influential Points
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Variation
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
Regression Toward the Mean
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
