Related Experiment Video
Updated: Feb 5, 2026

An Orthotopic Murine Model of Human Prostate Cancer Metastasis
Published on: September 18, 2013
A comparative evaluation of large language models for simplifying prostate cancer pathology reports: ChatGPT and
Haoyang Zeng1,2, Yangguang Yuan1, Xiang Wu3
1Department of Radiology, The Third Affiliated Hospital of Shenzhen University (Luohu People's Hospital), Shenzhen, China.
Objectives:
To evaluate the application value of three ChatGPT versions and Gemini in pathology report simplification tasks for prostate cancer.
Methods:
This retrospective study assessed GPT-3.5, GPT-4.0, GPT-4o, and Gemini on pathology reports from 228 prostate cancer patients across two institutions. Data were split into internal (center 1, n = 171) and external (center 2, n = 57) cohorts. Using specific prompts, models generated simplified texts. The evaluation of outputs included three main dimensions: (1) human scoring by patients, clinicians, and pathologists; (2) readability scores; and (3) BERT-based semantic similarity scores. Statistical comparisons employed paired t -tests or Wilcoxon signed-rank tests. Statistical consistency between raters was assessed using squared weighted kappa, intraclass correlation coefficient(3,1), and percent agreement, with 95% confidence intervals calculated for all metrics.
Results:
GPT-4o (Few-Shot) achieved the highest accuracy and comprehensiveness scores from pathologists, while Gemini demonstrated the best understandability. Patient and clinician understandability ratings were consistently high across models. Mean Reading Grade Level scores varied between internal and external datasets, with GPT-4o Few-Shot performing best overall. BERT-based semantic similarity scores demonstrated distinct trends across models, reflecting differences in text simplification strategies.
Conclusion:
LLMs adopt distinct trade-off strategies between simplifying pathology reports and preserving their structure and logic, influenced by prompt design and textual style. Their application shows potential to enhance patient comprehension and clinical communication. Future work should focus on domain-specific fine-tuning to ensure safe and reliable clinical integration.
More Related Videos
Related Concept Videos
Simplified Synchronous Machine Model
In this model, each generator is connected to a...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Self-Evaluation Maintenance Model

