Related Experiment Video
Updated: Jun 17, 2026

07:13
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Benchmarking reliability and calibration of LLMs for multi-cancer early detection test communication
Koki Takabatake1,2, Maria Sol Rosito2,3, Danielle Braun2,3
1Boston University Chobanian and Avedisian School of Medicine, Boston, MA, 02118, United States.
JAMIA Open
|June 16, 2026
Summary
Large Language Models (LLMs) show high accuracy in communicating multi-cancer early detection (MCED) test details, but struggle with numerical recall and confidence reliability, necessitating human oversight.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Oncology
Background:
- Multi-cancer early detection (MCED) tests are emerging tools in oncology.
- Accurate communication of MCED test characteristics is crucial for clinical adoption.
- Large Language Models (LLMs) show potential for summarizing complex medical information.
Purpose of the Study:
- To evaluate the reliability and confidence calibration of LLMs in communicating MCED test characteristics.
- To assess LLM performance in an expert-informed clinical context.
Main Methods:
- A question set was developed from 19 MCED papers, reviewed by experts, and comprised 94 MCQs and 13 FRQs.
- Five LLMs (ChatGPT, Claude, Gemini, Perplexity, OpenEvidence) answered the questions three times.
- Accuracy and confidence calibration were assessed using various metrics, including Brier scores and high-confidence error analysis.
Main Results:
- All evaluated LLMs demonstrated high overall accuracy in communicating MCED test characteristics.
- Gemini exhibited the highest performance across MCQ and FRQ formats.
- Systematic weaknesses were identified in numerical recall and confidence reliability despite high accuracy.
Conclusions:
- LLMs can assist in disseminating MCED information.
- Clinical implementation requires verification workflows and human oversight to mitigate identified weaknesses.
