Related Experiment Video
Updated: Jan 17, 2026

Interactive and Visualized Online Experimentation System for Engineering Education and Research
Published on: November 24, 2021
Comparative analysis of AI and expert evaluations in engineering design pedagogy
Tuğra Karademir Coşkun1, Esra Bozkurt Altan2
1Computer Education, Faculty of Education, Sinop, Turkiye.
Background:
Integrating engineering design processes into science education has become a significant priority in STEM instruction. However, many science teachers face difficulties incorporating these processes due to limited pedagogical expertise. Generative artificial intelligence (GAI) tools such as ChatGPT offer potential support mechanisms by evaluating lesson plans and providing formative feedback. This study investigates the reliability and validity of GAI evaluations compared to expert assessments.
Methods:
This mixed-methods study involved 43 science teachers who received professional development over four months to integrate engineering design into their lesson plans. A total of 52 lesson plans were evaluated using structured and unstructured prompts via ChatGPT 4.5, alongside evaluations by expert mentors. Quantitative data were analyzed using the Intraclass Correlation Coefficient (ICC) and Bland-Altman methods to assess inter-rater consistency. Qualitative data was analyzed through open and deductive coding to interpret differences in evaluation rationale.
Results:
Findings revealed high consistency between structured prompt AI evaluations and expert assessments (ICC = 0.708), while unstructured prompts showed low and non-significant agreement (ICC = 0.076). Qualitative analysis indicated that AI evaluations, particularly those using structured prompts, tend to be more positive and holistic, whereas experts offered more detailed and critical feedback. Differences were also observed in evaluating dcomponents like problem definition, testability, and interdisciplinary integration.
Conclusion:
Structured AI prompts offer reliable and valid evaluation results comparable to expert assessments and could serve as scalable tools in teacher support systems. However, unstructured prompts produce inconsistent outcomes and require refinement. The study highlights both the potential and limitations of using GAI tools for pedagogical evaluation in STEM education.
Related Concept Videos
Stereotype Content Model
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...