ChatGPT and academic integrity: a case study of assessment vulnerability in biological and environmental sciences
Danielle Hinchcliffe1, Chrysanthi Fergani1, Susanne R K Zajitschek1
1School of Biological and Environmental SciencesLiverpool John Moores UniversityLiverpoolUnited Kingdom.
Abstract:
The rise of generative artificial intelligence (AI) has accelerated content creation across sectors, including education and academia. While these tools can streamline workflows and support learning, their use by students to complete assignments poses risks to academic integrity. This case study quantitatively evaluated the vulnerability of existing coursework assessments and exams to cheating using generative AI. Specifically, we compared ChatGPT-4's performance to student grades on 131 assessments from 40 modules in the School of Biological and Environmental Sciences at a UK University. ChatGPT excelled in "exam-like" assessments, including coursework-associated tests, exams (particularly those involving multiple choice), and, at some levels of study, in essay-style assessments. Conversely, AI performance and hence vulnerability were low on assessments requiring in-class data collection, collaboration, or the creation of reports, posters, or presentations. Overall, ChatGPT-produced work was often superficial and vague but better than expected, posing risks to assessing learning in all fields within the biological sciences, including physiology. While generative AI holds promise as a tool to help students understand, organize, and structure, it raises concerns around misuse. Designing creative, individualized tasks that require genuine intellectual engagement is essential, with greater emphasis on ethical conduct and the importance of integrity. Synthesizing information across multiple biological scales and applying clinical or experimental scenarios is recommended to develop higher-order physiological reasoning in students. Continuous, collaborative monitoring of generative AI performance on assessments should become part of routine assessment design and evaluation, given the continuous development and improvement of these tools. In cases where basic knowledge testing remains necessary, reversion to in-person examinations is recommended to replace online testing.NEW & NOTEWORTHY This study reveals that current assessments in biological and environmental sciences are highly vulnerable to generative artificial intelligence (AI) misuse, which threatens assessment validity. ChatGPT-4 either outperformed or approximated students in exam-style testing and on some essay assignments but struggled with tasks involving data collection, collaboration, and presentations. The authors recommend redesigning assessments to require deeper reasoning, ethical engagement, and clinical or experimental application, alongside ongoing monitoring of AI capabilities and increased use of in-person exams.
Related Concept Videos
Methods to Assess Microbial Communities
Threats to Biodiversity
Biodeterioration
