Related Experiment Videos
A dataset of human and AI-generated rubric evaluations for formative programming assessment.
Pedro C Mendonça1,2, Filipe Quintal3,4, Karolina Baras3
1Faculty of Exact Sciences and Engineering, University of Madeira, 9000-082, Funchal, Portugal. pedro.mendonca@staff.uma.pt.
Scientific Data
|May 25, 2026
Summary
The HARMOGEN-R dataset aids research in automated rubric generation and evaluation using artificial intelligence. This dataset supports developing AI tools for faster, more frequent student feedback in education.
Area of Science:
- Artificial Intelligence in Education
- Educational Technology
- Computer Science Education
Background:
- Formative assessment is time-intensive due to manual rubric design and evaluation.
- Limited feedback frequency hinders student learning and development.
- Automated solutions are needed to streamline assessment processes.
Purpose of the Study:
- Introduce the HARMOGEN-R dataset for AI-driven rubric generation and evaluation.
- Facilitate research on supervised AI for educational assessment.
- Enable studies on rubric performance and feedback quality.
Main Methods:
- Collected 770 student responses from a Data Structures and Algorithms course.
- Used one human-created and four LLM-generated rubrics (structured and free-form).
- Performed parallel evaluations by human instructors and LLMs, recording reasoning traces and feedback reports.
- Implemented a variance-based quality control system for evaluation consistency.
Main Results:
- The dataset enables comparison of human vs. LLM rubric performance.
- Facilitates analysis of evaluator agreement and assessment reliability.
- Provides data on reasoning traces and feedback generation for AI assessment.
Conclusions:
- The HARMOGEN-R dataset is a valuable resource for advancing automated formative assessment.
- Supports research into AI-assisted rubric creation and evaluation under teacher supervision.
- Aims to improve the efficiency and quality of student feedback in programming courses.