Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: May 28, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

Diagnostic Performance and Confidence Calibration of Large Language Models for Bone Tumor Radiographs.

Sanjana Arun1, Eujung Park1, Katja Klosterman1

  • 1College of Medicine Phoenix, University of Arizona, Phoenix, AZ 85004, USA.

Diagnostics (Basel, Switzerland)
|May 27, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Functional and Cosmetic Outcomes of Pterional Craniotomy and Its Modifications: A Scoping Review.

Cureus·2026
Same author

Recommended Cardiometabolic Screening Guidelines for Unhoused Adults: A Street Medicine Needs Assessment.

Clinics and practice·2026
Same author

Cranioplasty Approaches and Outcomes in Low-Middle Income Countries: A Systematic Review.

The Journal of craniofacial surgery·2025
See all related articles

Large language models (LLMs) show promise in detecting bone tumors on radiographs but struggle with accurate classification, especially for challenging cases like Ewing sarcoma. Their current limitations, including false positives and overconfidence, suggest they are best used as assistive tools in musculoskeletal radiology.

Area of Science:

  • Radiology
  • Artificial Intelligence
  • Oncology

Background:

  • Large language models (LLMs) are increasingly used in medical imaging.
  • Their diagnostic accuracy and reliability in musculoskeletal radiology are not well-established.
  • This study focuses on LLMs' performance in detecting and classifying bone tumors on radiographs.

Purpose of the Study:

  • To evaluate the diagnostic performance of LLMs in detecting and classifying bone tumors on radiographs.
  • To assess the confidence calibration of LLMs in musculoskeletal radiology.
  • To compare the performance of three different LLMs: ChatGPT 5.3, X-ray Interpreter GPT-4.1, and X-ray Interpreter Gemini.

Main Methods:

  • A retrospective observational study analyzed 257 radiographs with confirmed bone tumor diagnoses from Radiopaedia.
Keywords:
artificial intelligencebone tumorsdiagnostic accuracylarge language modelsmedical imagingradiographs

Related Experiment Videos

Last Updated: May 28, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

  • Three LLMs evaluated images for abnormality detection, tumor detection, classification, and confidence.
  • Outcomes included diagnostic accuracy, false positive/negative rates, tumor hallucination, and confidence calibration.
  • Main Results:

    • LLMs demonstrated high abnormality detection rates, with Gemini showing the highest sensitivity.
    • Tumor detection was best for characteristic lesions like osteosarcoma and osteochondroma.
    • False negative rates varied (Gemini: 6.6%, ChatGPT: 24.8%, GPT-4.1: 29.9%); classification accuracy was poor for Ewing sarcoma.
    • False positive rates were highest for GPT-4.1 (40.7%).
    • All models exhibited confidence miscalibration, with overconfidence in incorrect predictions.

    Conclusions:

    • LLMs excel at detecting radiographic abnormalities but have limitations in tumor subtype classification, particularly for Ewing sarcoma.
    • High false positive/negative rates and overconfidence (especially GPT-4.1) restrict current clinical utility.
    • LLMs are better suited as adjunctive tools rather than standalone diagnostic systems in musculoskeletal radiology.