预测建模和队列数据分析为学生的成功和留守
1Computer Engineering and Computer Science Department, California State University, Long Beach, 90840, CA, USA.
Evaluation and program planning
|September 2, 2025
概括
学业成绩因学生的人口结构而异. 少数群体和符合Pell资格的学生面临挑战,但预测模型可以识别有风险的学生以提供有针对性的支持和提高留学时间.
科学领域:
- 高等教育研究
- 教育数据挖掘
- 学生成绩分析
背景情况:
- 学业成绩和留学率是高等教育的关键指标.
- 了解人口差异对于公平的学生支持至关重要.
- 预测模型为学生的成功提供了早期干预的机会.
研究的目的:
- 分析不同学生群体的学术表现差异.
- 确定影响学生成绩的关键因素 (GPA,学分积累).
- 开发和验证用于预测学生成绩的预测模型.
主要方法:
- 通过对23,000名新生进行数据驱动分析.
- 考量因素:平均成绩,学分积累,培训资格,少数民族身份,父母教育.
- 集群分析和深度学习用于预测建模.
主要成果:
- 观察到显著的差异:少数民族和符合Pell资格的学生积累的学分较少.
- 少数民族学生的平均成绩较低,平均成绩差异较大.
- 通过集群分析确定了三种不同的学术参与概况.
结论:
- 学生的不同表现需要不同的支持策略.
- 预测模型准确地预测了第二年学分积累和GPA.
- 研究结果为提高学生留学率和学术发展提供了可操作的见解.
相关概念视频
Steps in Outbreak Investigation
188
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
188
Longitudinal Research
12.4K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.4K
Statistical Methods for Analyzing Epidemiological Data
530
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
530
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Longitudinal Studies
231
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
231
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


