Related Experiment Video
Updated: Jun 14, 2026

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Evaluating the Pediatric Behavior Guidance of Students Based on Actual Clinical Transcripts Scored by Faculty and
Ishreen Kaur Dhillon1, Gabriel Keng Yan Lee1, Shijia Hu1
1Faculty of Dentistry, National University of Singapore, 9 Lower Kent Ridge Road, Singapore, 119085, Singapore, 65 67727757, 65 67785742.
Background:
Personalized feedback improves the clinical pediatric behavior guidance performance of students but is prohibitively time-consuming to provide. Large language models (LLMs) can automate the process of evaluating clinical sessions but are limited to text-only input and consistency issues.
Objective:
This study compared the use of text-only transcripts against the use of video recordings for evaluating the clinical behavior guidance performance of dental students. Additionally, the consistency and accuracy of LLMs in evaluating the transcripts were compared against a human assessor.
Methods:
This study was conducted by using 40 video-recorded clinical encounters involving final-year dental students who were managing patients aged between 4 and 12 years at the Faculty of Dentistry, National University of Singapore. The videos were scored by using a previously validated pediatric behavior guidance scale. Clinical encounters were transcribed verbatim and scored by a study member using a modified version of the scale (nonverbal components removed). The time taken to rate the transcripts was recorded. Video scores were compared with transcript scores. Both the free-to-use version and the paid version of the ChatGPT LLM were also used to score the transcripts; consistency was evaluated and compared against the human assessor.
Results:
The average time taken to rate the transcripts (mean 12, range 3-25 min) was significantly (P<.001) lower than the average video length (mean 73, range 37-120 min). Comparing transcript scores with video scores resulted in a consistency intraclass correlation coefficient of 0.830 (95% CI 0.679-0.910; P<.001), demonstrating good reliability. Comparing transcript scores with the free-to-use LLM's and paid LLM's scores yielded an absolute agreement intraclass correlation coefficient of 0.729 (95% CI 0.475-0.859; P<.001) and 0.670 (95% CI 0.377-0.825; P<.001), respectively, demonstrating moderate agreement. The LLMs were inconsistent, producing variable scores with the same prompt. The free-to-use and paid versions produced the same score for all 3 runs in only 7 (18%) and 4 (10%) of the 40 clinical encounters, respectively.
Conclusions:
Using transcripts to evaluate students' clinical behavior guidance was time-saving for faculty, demonstrated good agreement with video-based evaluation, and could improve clinical teaching. Although LLMs can automate the task, improvements are needed to improve their consistency and accuracy.