Related Experiment Video
Updated: Jun 11, 2025

Improving Student Outcomes with an Adaptable Molecular Cloning Course-Based Undergraduate Research Experience
Published on: November 15, 2024
A comparison of human, GPT-3.5, and GPT-4 performance in a university-level coding course
Will Yeadon1, Alex Peach2, Craig Testrow2
1Department of Physics, Durham University, Durham, DH1 3LB, UK. will.yeadon@durham.ac.uk.
Abstract:
This study evaluates the performance of ChatGPT variants, GPT-3.5 and GPT-4, both with and without prompt engineering, against solely student work and a mixed category containing both student and GPT-4 contributions in university-level physics coding assignments using the Python language. Comparing 50 student submissions to 50 AI-generated submissions across different categories, and marked blindly by three independent markers, we amassed data points. Students averaged 91.9% (SE:0.4), surpassing the highest performing AI submission category, GPT-4 with prompt engineering, which scored 81.1% (SE:0.8)-a statistically significant difference (p = ). Prompt engineering significantly improved scores for both GPT-4 (p = ) and GPT-3.5 (p = ). Additionally, the blinded markers were tasked with guessing the authorship of the submissions on a four-point Likert scale from 'Definitely AI' to 'Definitely Human'. They accurately identified the authorship, with 92.1% of the work categorized as 'Definitely Human' being human-authored. Simplifying this to a binary 'AI' or 'Human' categorization resulted in an average accuracy rate of 85.3%. These findings suggest that while AI-generated work closely approaches the quality of university students' work, it often remains detectable by human evaluators.
More Related Videos
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
06:09P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
Published on: September 8, 2023
Related Concept Videos
Comparing Experimental Results: Student's t-Test
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Improving Translational Accuracy
Reliability and Validity
Evolutionary Relationships through Genome Comparisons
Machines: Problem Solving II