Related Experiment Video
Updated: Sep 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Pilot Validation of a Large Language Model Facilitator for Peer-to-Peer Learning
Simon Hollingsworth1, Tom McIntyre1, Sami Ahmed1
1Discipline of Surgery, University of Dublin, Trinity College, Dublin, Ireland; Department of Surgery, Tallaght University Hospital, Dublin, Ireland.
Objective:
Active learning strategies such as peer-to-peer learning (P2PL) are being adopted by medical schools worldwide. P2PL involves students of the same level teaching each other without direct instructor involvement. Maintaining accuracy and consistency of information presented during P2PL necessitates employment of expert facilitators. However, additional challenges exist, such as increased costs, differences in learning opportunities, and loss of the key benefits of P2PL due to inter-facilitator differences. We aim to address these challenges by investigating three LLMs as facilitators for P2PL in surgical education amongst final year medical students.
Design:
65 final year medical students were recruited. P2PL presentations were scored in real-time by surgical tutors and qualitative feedback provided. P2PL presentations were recorded, transcribed, anonymised, and uploaded to three LLMs, ChatGPT, Gemini, and Claude. LLMs were prompted to score student presentations across ten metrics; identify low-confidence statements and factual errors; and provide qualitative feedback. Agreement rates between LLM and surgical tutor scoring and inter-rater reliability (IRR) was calculated.
Results:
There was no statistically significant difference in mean score between tutors and Chat GPT and Claude (p>0.05), but a statistically significant increased level of scoring seen with Gemini (p<0.0001). IRR was assessed using intraclass correlation co-efficient (ICC). Poor reliability was seen for mean student score (ICC=0.39), moderate reliability for four scoring metrics (ICC≥0.50-<0.70) and poor reliability for six metrics (ICC<0.50). Analysis of LLM ability to identify factual errors and low confidence statements revealed poor IRR (Fleiss κ<0.20).
Conclusions:
This study validates the technical feasibility of LLMs in P2PL within undergraduate surgical education. However, our findings suggest LLMs are not currently capable of replacing human educators, rather they should be used as valuable tools to improve the educational experience for students. This study acts as a framework for further research in this area, ultimately paving the way for the development of a personalised LLM tutor for students.
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Purposive Learning
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Observational Learning
Language and Cognition
Learning Disabilities
Dyslexia
Dyslexia is a...