Related Experiment Video
Updated: Sep 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using Large Language Models to Analyze Symptom Discussions and Recommendations in Clinical Encounters
Anny T H R Fenton1, Natasha Charewycz1, Zarwah Kanwal1
1Department of Medical Oncology, Dana-Farber Cancer Institute, Boston, Massachusetts, USA.
Large language models (LLMs) can accurately analyze patient-provider interactions, showing strong agreement with human coders. This technology offers a feasible tool for improving healthcare quality and communication by analyzing clinical encounters.
Area of Science:
- Health Informatics
- Artificial Intelligence in Medicine
- Clinical Communication Research
Background:
- Analyzing patient-provider interactions is crucial for assessing care quality but is often hindered by time and methodological challenges.
- Large language models (LLMs) present a potential solution for analyzing these interactions, but their accuracy needs validation.
- Existing methods for evaluating clinical communication are resource-intensive, limiting their widespread application.
Purpose of the Study:
- To evaluate the accuracy and reliability of a large language model (LLM) in analyzing patient-provider communication within clinical encounter transcripts.
- To compare the coding performance of an LLM against human coders on key aspects of symptom discussions.
- To determine the feasibility of using LLMs as a research tool for analyzing patient-provider interactions.
Main Methods:
- A large language model (GPT-4) was used to code 236 potential symptom discussions from 92 cancer patient clinical transcripts.
- Human coders independently analyzed the same transcripts to identify symptom discussion, initiation, and recommendations.
- Cohen's kappa (κ) was calculated to measure interrater agreement between the LLM and human coders.
Main Results:
- The LLM demonstrated strong to moderate interrater reliability with human coders across all measures.
- Highest agreement was observed for symptom discussion (κ = 0.89), followed by initiation (κ = 0.82) and recommendations (κ = 0.78).
- Disagreements regarding recommendations occurred in 16% of cases, categorized into nine distinct types.
Conclusions:
- LLMs show comparable analytical abilities to humans in evaluating patient-provider interactions from clinical transcripts.
- The use of LLMs can significantly enhance the feasibility of analyzing patient-provider communication for research purposes.
- LLM-driven analysis holds potential for broader applications in assessing care quality, identifying inequities, and improving communication.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Formulating and Validating Nursing Diagnosis I
There are thirteen domains...
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...