Related Experiment Videos
High-Stakes AI Feedback in Clinical Assessment: A Comparative Evaluation of GPT-4o and Claude 4 Feedback Fidelity
Thomas Kropmans1, Oleh Bilokrylyi1, Dmytro Predchyshyn2
1Qpercom Ltd., Galway, Ireland.
Abstract:
This study compares OpenAI's GPT-4o and Anthropic's Claude 4 in the generation of formative and summative feedback in Objective Structured Clinical Examinations (OSCEs) within Qpercom's assessment platform. A stratified sample of 51 anonymized student records was analyzed, comparing examiner-facing (pre-verification/preview) and student-facing (portfolio) feedback across both models. While both systems delivered actionable suggestions, Claude 4 consistently outperformed GPT-4o in alignment with examiner data, absence of hallucinations, and preservation of critical learning points-especially for underperforming and mid-performing students. This evidence-based evaluation recommends Claude 4 as the safer and more effective AI solution for high-stakes educational settings.
Related Concept Videos
Effects of feedback
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
Feedback control systems
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Sources of Self-Esteem II: Performance Feedback