Benchmarking reliability and calibration of LLMs for multi-cancer early detection test communication

Koki Takabatake1,2, Maria Sol Rosito2,3, Danielle Braun2,3

  • 1Boston University Chobanian and Avedisian School of Medicine, Boston, MA, 02118, United States.

JAMIA Open
|June 16, 2026
PubMed
Summary

Large Language Models (LLMs) show high accuracy in communicating multi-cancer early detection (MCED) test details, but struggle with numerical recall and confidence reliability, necessitating human oversight.

Related Concept Videos