Related Experiment Video
Updated: Jun 7, 2026

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
A machine learning framework with a public synthetic data set for early student profiling and outcome prediction
1Department of Artificial Intelligence and Data Science, Faculty of Engineering and Technology, Parul University, Vadodara, Gujarat, India. sanjay.agal32685@paruluniversity.ac.in.
Scientific Reports
|June 5, 2026
Summary
This study introduces a synthetic dataset and machine learning framework for predicting student academic success. Ensemble methods and ordinal neural networks show strong performance, enabling data-driven interventions.
Area of Science:
- Educational Data Mining
- Machine Learning in Education
- Student Performance Prediction
Background:
- Accurate student outcome prediction is crucial for timely interventions in higher education.
- Privacy regulations limit access to real student data, hindering educational data mining research.
- A novel, public synthetic dataset of over 100,000 student records is introduced to overcome data scarcity.
Purpose of the Study:
- To develop and validate a comprehensive machine learning framework for early student profiling and outcome prediction.
- To create and release a public synthetic dataset suitable for machine learning research in educational contexts.
- To establish a benchmark for student outcome prediction and provide resources for reproducible research.
Main Methods:
- Development of a modular machine learning framework including preprocessing, feature engineering, and diverse algorithms (baselines, ensembles, neural networks).
- Definition of two prediction tasks: ordinal classification of student academic level and nominal classification of division allotment.
- Rigorous cross-validation and multi-faceted evaluation using accuracy, F1 scores, and ordinal-specific measures.
Main Results:
- Ensemble methods, particularly LightGBM, significantly outperformed simpler models.
- An ordinal neural network achieved the highest performance (Macro F1: 0.846, Quadratic Weighted Kappa: 0.892) for student level prediction.
- Prior academic achievement (Class 12 percentage) was the most dominant predictor, with socio-economic and behavioral factors offering secondary contributions.
Conclusions:
- The developed framework and synthetic dataset provide essential resources for advancing reproducible educational data mining research.
- Explicit ordinal modeling enhances prediction accuracy for student academic levels.
- The findings support the development of equitable, data-driven student support systems by identifying key predictive factors and enabling targeted interventions.