Related Experiment Video
Updated: Sep 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using Large Language Models to Analyze Symptom Discussions and Recommendations in Clinical Encounters
Anny T H R Fenton1, Natasha Charewycz1, Zarwah Kanwal1
1Department of Medical Oncology, Dana-Farber Cancer Institute, Boston, Massachusetts, USA.
Abstract:
Background: Patient-provider interactions could inform care quality and communication but are rarely leveraged because collecting and analyzing them is both time-consuming and methodologically complex. The growing availability of large language models (LLMs) makes these analyses more feasible, though their accuracy remains uncertain. Objectives: Assess an LLM's ability to analyze patient-provider interactions. Design: Compare a human's and an LLM's codings of clinical encounter transcripts. Setting/Subjects: Two hundred and thirty-six potential symptom discussions from transcripts of clinical encounters with 92 patients living with cancer in the mid-Atlantic United States. Transcripts were analyzed by GPT4DFCI in our hospital's Health Insurance Portability and Accountability Act compliant infrastructure instance of GPT-4 (OpenAI). Measurements: Human and an LLM-coded transcripts to determine whether a patient's reported symptom(s) were discussed, who initiated the discussion, and any resulting recommendation. We calculated Cohen's κ to assess interrater agreement between the LLM and human and qualitatively classified disagreements about recommendations. Results: Interrater reliability indicated "strong" and "moderate" agreement levels across measures: Agreement was strongest for whether the symptom was discussed (k = 0.89), followed by who initiated the discussion (k = 0.82), and the recommendation provided (k = 0.78). The human and LLM disagreed on the presence and/or content of the recommendation in 16% of potential discussions, which we categorized into nine types of disagreements. Conclusions: Our results suggest that LLMs' abilities to analyze clinical encounters are equivalent to humans. Thus, using LLMs as a research tool may make it more feasible to analyze patient-provider interactions, which could have broader implications for assessing and improving care quality, care inequities, and provider communication.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Formulating and Validating Nursing Diagnosis I
There are thirteen domains...
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...