Related Experiment Video
Updated: Sep 8, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
680
Scoring Physician Risk Communication in Prostate Cancer Using Large Language Models
Guillermo Lopez-Garcia1, Dongfang Xu1, Michael Luu2
1Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, USA.
Medrxiv : the Preprint Server for Health Sciences
|August 20, 2025
Summary
This study introduces an AI framework using large language models to automatically assess physician risk communication quality in prostate cancer care. This scalable solution improves upon manual evaluation, enhancing shared decision-making.
Area of Science:
- Oncology
- Medical Informatics
- Artificial Intelligence
Background:
- Effective risk communication is crucial for shared decision-making in prostate cancer treatment.
- Current manual evaluation of physician communication quality is time-consuming and not scalable.
- Variability in communicating treatment tradeoffs impacts patient understanding and choices.
Purpose of the Study:
- To develop and validate a structured, rubric-based framework using large language models (LLMs) for automated scoring of risk communication quality.
- To assess the performance of LLMs, specifically GPT-4o, in evaluating physician communication in prostate cancer consultations.
- To establish a scalable method for analyzing physician-patient communication in oncology.
Main Methods:
- A rubric-based framework was developed to score physician communication on five key domains: cancer prognosis, life expectancy, and three treatment side effects.
- 487 physician-spoken sentences from 20 clinical visit transcripts were annotated using a 0-5 scoring rubric for precision and patient-specificity.
- The task was modeled as multiclass classification, evaluating fine-tuned transformers and GPT-4o with rubric-based and chain-of-thought (CoT) prompting, including few-shot learning.
Main Results:
- The best performing approach, combining rubric-based CoT prompting with few-shot learning, achieved micro-averaged F1 scores between 85.0% and 92.0% across the evaluated domains.
- This AI-driven method outperformed traditional supervised baseline models.
- The developed framework demonstrated performance comparable to human inter-annotator agreement.
Conclusions:
- Large language models, particularly GPT-4o with rubric-based CoT prompting and few-shot learning, offer a scalable and accurate method for evaluating physician risk communication in prostate cancer care.
- This AI-driven approach provides a foundation for automated assessment of physician-patient communication, potentially improving shared decision-making.
- The findings have implications for quality improvement initiatives and training in oncology and other medical fields.
Keywords:
Prostate cancerartificial intelligencelarge language modelsnatural language processingphysician-patient communicationrisk communicationshared decision-making
