Related Experiment Video
Updated: Sep 9, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Predictive modeling and cohort data analytics for student success and retention
1Computer Engineering and Computer Science Department, California State University, Long Beach, 90840, CA, USA.
Abstract:
This study presents a data-driven analysis of academic performance, demographic disparities, and predictive modeling among more than 23,000 first-time freshmen at a US public University. We examine multiple factors influencing student outcomes, including GPA, credit accumulation, unit workload, Pell Grant eligibility, minority status, and parent education levels. Our analysis reveals several statistically significant disparities: non-minority students earn more units than minority students in their first two years, and Pell-eligible students accumulate fewer credits than their non-eligible peers. First-generation college students also exhibit lower credit accumulation compared to peers. GPA distributions show that minority students have a lower average GPA compared to non-minority students, with broader variation. Clustering analysis identifies three distinct academic engagement profiles based on GPA and unit load, highlighting heterogeneous performance patterns and the need for differentiated support. We develop and tune predictive models to forecast sophomore credit accumulation and GPA, achieving strong performance using deep learning. These models enable proactive risk identification and support strategic interventions. Our findings set the stage for actionable insights for institutional decision-makers aiming to enhance student retention, success, and academic momentum.
Related Concept Videos
Steps in Outbreak Investigation
Longitudinal Research
Statistical Methods for Analyzing Epidemiological Data
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Longitudinal Studies
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

