Related Experiment Video
Updated: Apr 8, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Exploring the feasibility of inferring prostate cancer pathological grade from multiparametric MRI text reports using
Yao Niu1, Liting Shen2, Jin Liu3
1Department of Interventional Medicine, Beijing Chaoyang Hospital, Capital Medical University, Beijing, China.
Objectives:
This study conducted a natural language processing feasibility analysis aimed at comparing four large language models (LLMs) in terms of (a) reproducibility and (b) predictive accuracy for International Society of Urological Pathology Grade Groups (ISUP GGs) based on structured text reports from prostate multiparametric magnetic resonance imaging (mpMRI).
Methods:
The study first used LLMs to perform the initial round of ISUP GGs predictions based solely on the mpMRI text reports. This was followed by a second round of predictions that incorporated clinical information. Each prediction round was repeated three times to assess consistency. Three radiologists independently completed the first two rounds of ISUP GG predictions and then performed a third round of assessment after reviewing the LLMs' predictions. The study recorded the response times.
Results:
The study included 150 patients (median age, 69 years). Statistically significant differences were observed among different ISUP GGs in terms of age, PSA levels, prostate volume, PSA density, and PI-RADS scores. The four LLMs demonstrated good to excellent reproducibility (Kappa 0.671-0.861). ChatGPT-4.1 had the shortest response time (0.95-17.19 s). Furthermore, the study found that the accuracy of the LLMs (32.7-50.0%) was significantly lower than that of senior radiologist (72.7-76.0%) and intermediate-level radiologist (66.0-68.7%), but was comparable to that of junior radiologist (59.3-65.3%).
Conclusion:
General-purpose LLMs demonstrate excellent reproducibility. While ChatGPT-4.1 outperforms other LLMs in ISUP GGs prediction and response time, its predictive accuracy remains inferior to that of intermediate and senior radiologists. Therefore, specific fine-tuning of this technology is necessary before general-purpose LLMs can be applied in clinical practice.
More Related Videos
08:05Detection and Isolation of Cancer in Prostate Biopsies Using Stimulated Raman Histology and Artificial Intelligence
Published on: June 10, 2025
06:08A Cognitive Fusion-guided Prostate Biopsy Using Multiparametric Magnetic Resonance Imaging and Transrectal Ultrasound
Published on: March 21, 2025