Related Experiment Video
Updated: May 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The In-depth Comparative Analysis of Four Large Language AI Models for Risk Assessment and Information Retrieval from
Lun-Hsiang Yuan1,2, Shi-Wei Huang2,3, Dean Chou1,4,5,6
1Department of Biomedical Engineering, National Cheng-Kung University, Tainan, Taiwan.
Four large language models (LLMs) were evaluated for information retrieval and risk assessment in prostate cancer (PC) reports. ChatGPT-4-turbo showed the highest accuracy in risk assessment, indicating potential for clinical decision support.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Oncology
- Clinical Decision Support Systems
Background:
- Accurate information retrieval and risk assessment from multi-modality reports are crucial for effective prostate cancer (PC) treatment.
- Large language models (LLMs) offer potential for automating these complex tasks.
Purpose of the Study:
- To evaluate the performance of four general-purpose LLMs in information retrieval (IR) and risk assessment (RA) tasks using simulated clinical data for stage IV PC.
- To compare the capabilities of ChatGPT-4-turbo, Claude-3-opus, Gemini-Pro-1.0, and ChatGPT-3.5-turbo in handling diverse clinical data.
Main Methods:
- Simulated multi-modality reports (CT, MRI, bone scans, pathology) for 350 stage IV PC patients were used.
- Four LLMs were assessed on seven IR tasks (including TNM staging) and three RA tasks (LATITUDE, CHAARTED, TwNHI).
- Zero-shot chain-of-thought prompting and ensemble voting methods were employed, with consensus from three adjudicators serving as the gold standard.
Main Results:
- LLMs demonstrated high accuracy (87.4%-94.2%) and consistency (ICC>0.8) in TNM staging.
- Significant differences were observed in RA performance, with ChatGPT-4-turbo outperforming others (accuracy 90.1%-91.6%).
- Ensemble voting consistently improved accuracy and reliability compared to single queries.
Conclusions:
- ChatGPT-4-turbo exhibited satisfactory performance in both RA and IR for stage IV PC, suggesting its utility in clinical decision support.
- Despite promising results, the potential for misinterpretation necessitates caution and further validation in diverse oncological contexts.
More Related Videos
06:08A Cognitive Fusion-guided Prostate Biopsy Using Multiparametric Magnetic Resonance Imaging and Transrectal Ultrasound
Published on: March 21, 2025
13:19Microarray-based Identification of Individual HERV Loci Expression: Application to Biomarker Discovery in Prostate Cancer
Published on: November 2, 2013