Related Experiment Videos
A comparative evaluation of LLM-generated and clinician-written psychotherapy session summaries
Liran Keren1, Ayal Klein2, Yael Bar-Shachar1
1Department of Psychology, Bar-Ilan University, Ramat-Gan, Israel.
Objective:
Psychotherapy session summaries are crucial to the continuity of care, supervision, and treatment planning; however, they are time-consuming and contribute to the documentation burden. Large language models (LLMs) may help streamline this workflow, but their adequacy must be established. We evaluated whether state-of-the-art LLMs can generate clinically meaningful summaries comparable to those of human clinicians.
Method:
We compared 1,008 segment-level summaries (504 clinician-written and 504 LLM-generated) drawn from 108 five-minute segments across 36 sessions. The summaries adhered to a structured template capturing core therapeutic components. The LLM summaries were generated using a multi-step prompting pipeline. Blinded-to-source judges evaluated the summaries on factual consistency, salient meaning preservation, and clinical value.
Results:
Clinicians-written summaries were rated higher overall than LLM-generated summaries across the evaluation dimensions. Component-level analyses revealed exceptions: the LLM was superior in preserving central themes, and no differences were found regarding the clinical value of the therapeutic relationship and intrapersonal functioning. Overall, both sources were rated favorably.
Conclusion:
LLMs can produce clinically usable session summaries, but they do not fully capture the clinical subtleties or depth reflected in clinician-written summaries. These findings support clinician-supervised human-AI workflows where LLMs generate drafts that are reviewed and verified, rather than replacing clinicians.
Clinical Or Methodological Significance Of This Article:
While producing high-quality session summaries is clinically vital, its time-intensive nature creates a substantial documentation burden that emerging generative AI technologies could potentially alleviate. This study introduces a clinically informed, multi-step prompting pipeline and provides the first direct comparison of LLM-generated and clinician-written summaries across core therapeutic components using the transtheoretical MIND framework. The findings demonstrate that while LLMs capture structured session content, they do not yet match clinicians' nuanced understanding, supporting a collaborative "draft-and-review" paradigm where AI assists rather than replaces clinical reflection and judgment.