Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 4, 2026

Coordinate Mapping of Hyolaryngeal Mechanics in Swallowing
14:13

Coordinate Mapping of Hyolaryngeal Mechanics in Swallowing

Published on: May 6, 2014

Confidence Measurement Metrics in Multimodal Large Language Models for Ultrasound-Based Radiology Cases: Comparative

Taewon Han1, Jaeseung Shin1, Jeong Hyun Lee1

  • 1Department of Radiology, Samsung Medical Center, 81 Irwon-ro, Irwon-dong - Gangnam-gu, Seoul, 06351, Republic of Korea, 82 10-8714-7650.

Journal of Medical Internet Research
|June 2, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Correction: Confidence Measurement Metrics in Multimodal Large Language Models for Ultrasound-Based Radiology Cases: Comparative Evaluation Study of Self-Reported, Consistency-Based, and Hybrid Methods.

Journal of medical Internet research·2026
Same author

Agentic Artificial Intelligence for the Automated Generation of Accurate Summary Podcasts of Radiology Research Papers.

Korean journal of radiology·2026
Same author

Evaluating Accuracy of LI-RADS Nonradiation Treatment Response Algorithm v2024 and Ancillary Features at Hepatobiliary MRI versus CT.

Radiology·2026
Same author

LLM Label Noise and the Established Framework of Imperfect Reference Standard Bias.

Radiology. Artificial intelligence·2026
Same author

Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM): 2025 Updates.

Korean journal of radiology·2025
Same author

Academic Journal Podcast and Generative Artificial Intelligence: Introducing KJR SummaryCast.

Korean journal of radiology·2025

A hybrid confidence metric, the Top Weighted Score, reliably indicates diagnostic confidence in large language models (LLMs) interpreting radiology cases. This metric outperforms self-reported or consistency-based methods, supporting safer clinical deployment of AI in healthcare.

Area of Science:

  • Artificial Intelligence in Medical Imaging
  • Machine Learning for Diagnostic Support
  • Healthcare AI Confidence Assessment

Background:

  • Large language models (LLMs) require robust confidence assessment for safe integration into healthcare.
  • Existing methods for quantifying LLM confidence are insufficient for clinical applications.
  • Multimodal LLMs show promise in interpreting complex medical data like radiology images.

Purpose of the Study:

  • To evaluate confidence metrics for multimodal LLMs in ultrasound-based radiology.
  • To compare self-reported, consistency-based, and hybrid confidence assessment methods.
  • To identify reliable metrics for gauging LLM diagnostic confidence.

Main Methods:

  • Evaluated four multimodal LLMs (GPT-5, Claude-4.5-Sonnet, Gemini-3-Pro, GPT-4o) on 94 radiology cases.
Keywords:
AILLMsartificial intelligencediagnostic confidencelarge language modelsmedical informaticsradiology

More Related Videos

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
07:13

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities

Published on: October 27, 2023

Related Experiment Videos

Last Updated: Jun 4, 2026

Coordinate Mapping of Hyolaryngeal Mechanics in Swallowing
14:13

Coordinate Mapping of Hyolaryngeal Mechanics in Swallowing

Published on: May 6, 2014

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
07:13

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities

Published on: October 27, 2023

  • Assessed confidence using self-reported, consistency-based (relative entropy, majority-vote), and hybrid (Top Weighted Score) metrics.
  • Analyzed diagnostic accuracy, metric correlation with accuracy, model calibration, and resource utilization.
  • Main Results:

    • Gemini-3-Pro achieved the highest diagnostic accuracy (74.47%), exceeding median human performance.
    • The Top Weighted Score demonstrated statistically significant correlations with accuracy across all models.
    • Top Weighted Score exhibited superior discriminative ability and calibration, particularly for GPT-5 and Claude-4.5-Sonnet.

    Conclusions:

    • Hybrid confidence metrics, like the Top Weighted Score, are more reliable indicators of diagnostic confidence than individual methods.
    • Integrative confidence estimation approaches are crucial for safe clinical deployment of LLMs in radiology.
    • Further external validation is necessary before widespread clinical application of these LLM confidence metrics.