机器学习预测高中教育学早在小学结束时就开始了
Maria Psyridou1, Fabi Prezja2, Minna Torppa3
1Department of Psychology, University of Jyväskylä, 40014, Jyväskylä, Finland. maria.m.psyridou@jyu.fi.
Scientific reports
|June 5, 2024
概括
这项研究使用了对13年数据集的机器学习来预测放学率. 模型显示到9年级的准确性得到改善,为早期学生风险识别提供了潜力.
科学领域:
- 教育研究教育研究
- 机器学习应用程序 机器学习应用程序
- 社会发展社会发展.
背景情况:
- 学校学对个人和社会产生重大影响,阻碍了减贫和经济增长.
- 之前关于脱学预测的机器学习研究经常使用有限的短期数据.
- 需要一个长期的,全面的方法来理解和减轻学因素.
研究的目的:
- 开发和评估机器学习模型,利用13年的纵向数据集预测学校学情况.
- 评估学术,认知,行为和福祉因素对留学生的预测能力.
- 探索这些模型在支持主动教育干预方面的潜力.
主要方法:
- 利用了13年的纵向数据集 (幼儿园到九年级),包括学生的学术,认知,动机,行为和福祉数据.
- 开发并验证机器学习模型用于学分类.
- 使用曲线下的面积 (AUC) 度量在不同时间点 (到6级和9级) 评估模型性能.
主要成果:
- 机器学习模型使用高达6年级的数据实现了0.61的平均AUC.
- 当将数据纳入到第9级时,模型性能提高到0.65的AUC.
- 这些发现表明,纵向数据在提高学预测准确度方面的潜力.
结论:
- 纵向数据显著提高了机器学习模型对学校学的预测能力.
- 这些模型显示,积极识别有风险的学生,支持教育工作者是有希望的.
- 建议通过相关性和因果分析进行进一步的研究,以完善留学生策略.
相关概念视频
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Correlations
32.8K
Correlation means that there is a relationship between two or more variables (such as ice cream consumption and crime), but this relationship does not necessarily imply cause and effect. When two variables are correlated, it simply means that as one variable changes, so does the other. We can measure correlation by calculating a statistic known as a correlation coefficient. A correlation coefficient is a number from -1 to +1 that indicates the strength and direction of the relationship between...
32.8K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Applications of Life Tables
59
Life tables are versatile across various fields, providing a quantitative basis for analyzing mortality and survival rates. Whether used by demographers, actuaries, epidemiologists, or sociologists, life tables offer valuable insights into the dynamics of life and death, facilitating informed decisions in public health, insurance, conservation, and beyond. Their broad applicability highlights the interconnectedness of demographic data with practical outcomes in everyday life and strategic...
59


