使用机器学习来预测影响学业表现的因素:大学生在学术试用期间的情况
Lamees Al-Alawi1, Jamil Al Shaqsi1, Ali Tarhini1
1Department of Information Systems, College of Economics and Political Science, Sultan Qaboos University, P.O. Box 20, PC 123 Muscat, Oman.
概括
机器学习确定了影响大学生学业绩的关键因素. 学习时间和以前的中学成绩是表现不佳学生最显著的负面预测因素.
科学领域:
- 教育数据挖掘教育数据挖掘
- 高等教育中的机器学习
- 学生表现分析 学生表现分析
背景情况:
- 学术试用影响了许多大学生.
- 识别影响低绩效的因素对于干预至关重要.
- 之前的研究已经探索了各种有助于因素.
研究的目的:
- 应用监督机器学习算法,以识别影响试用大学生学业绩的负面因素.
- 用数据库中的知识发现 (KDD) 方法来进行数据分析.
- 为了比较不同的机器学习算法在预测学术表现不佳方面的有效性.
主要方法:
- 分析了来自阿曼一所公立大学的6514名大学生 (2009-2019) 的样本.
- 信息获取 (InfoGain) 算法用于特征选择.
- 组合方法 (Logit Boost, Vote, Bagging) 被使用并使用标准指标和10倍交叉验证进行评估.
主要成果:
- 研究时间和以前的中学成绩被确定为对学业成绩产生负面影响的主要因素.
- 性别,预计的毕业年,队列和学术专业化也对学生处于试用期做出了重大贡献.
- 使用的机器学习模型在识别这些影响因素方面表现强.
结论:
- 机器学习有效地识别了与高等教育学术表现不佳相关的关键因素.
- 大学的干预措施应考虑学习的持续时间和以前的学术历史.
- 进一步的研究可以完善个性化学生支持的预测模型.
相关概念视频
Reliability and Validity
12.8K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.8K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Correlations
33.4K
Correlation means that there is a relationship between two or more variables (such as ice cream consumption and crime), but this relationship does not necessarily imply cause and effect. When two variables are correlated, it simply means that as one variable changes, so does the other. We can measure correlation by calculating a statistic known as a correlation coefficient. A correlation coefficient is a number from -1 to +1 that indicates the strength and direction of the relationship between...
33.4K
Regression Analysis
5.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.8K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Outliers and Influential Points
4.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.1K


