应用机器学习技术的性能作为预测学校学的支持
Auria Lucia Jiménez-Gutiérrez1, Cinthya Ivonne Mota-Hernández2, Efrén Mezura-Montes3
1Centro Universitario de los Lagos, Universidad de Guadalajara, Enrique Díaz de León 1144, Paseos de la Montaña, 47460, Lagos de Moreno, Jalisco, Mexico. auria.jimenez@academicos.udg.mx.
Scientific reports
|February 17, 2024
概括
这项研究开发了一种机器学习模型,用于预测墨西哥的学业学率,达到高准确度. 该模型利用人口普查数据和各种算法来识别中学和高等教育中面临风险的学生.
科学领域:
- 教育数据挖掘教育数据挖掘
- 机器学习在教育中的应用
背景情况:
- 学业学率在中学和高等教育中构成了重大挑战.
- 预测建模可以帮助制定早期干预策略,以减少学业学率.
研究的目的:
- 设计和验证一种机器学习模型,用于预测墨西哥教育机构中放学者的情况.
- 为了实现至少90%的模型可靠性,用于学预测.
主要方法:
- 利用了墨西哥国家统计和地理研究所 (2010,2015,2020年人口普查) 的开放数据.
- 在数据同质化和相关性分析后,选择了20个相关变量,结果为1,080,782条记录.
- 应用监督学习技术,包括人工神经网络,支持矢量机器,线性/拉索回归,贝叶斯优化和随机森林.
主要成果:
- 人工神经网络和支持矢量机器的可靠性超过99%.
- 随机森林实现了91%的可靠性,达到研究的目标.
- 开发的模型有效地利用历史的人口和住房数据预测学校学.
结论:
- 机器学习模型,特别是ANN和SVM,在预测放学率方面非常有效.
- 该研究为教育机构提供了一种可靠的工具,可以主动解决学生退学问题.
- 来自人口普查信息的数据驱动的见解可以为学生留学提供有针对性的干预措施.
相关概念视频
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Applications of Life Tables
64
Life tables are versatile across various fields, providing a quantitative basis for analyzing mortality and survival rates. Whether used by demographers, actuaries, epidemiologists, or sociologists, life tables offer valuable insights into the dynamics of life and death, facilitating informed decisions in public health, insurance, conservation, and beyond. Their broad applicability highlights the interconnectedness of demographic data with practical outcomes in everyday life and strategic...
64
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Survival Tree
85
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
85
Microsoft Excel: Regression Analysis
605
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
605
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K


