Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Functional Outcome and Hip Survival Rate in Traumatic Femoral Head Fractures.

Journal of clinical medicine·2026
Same author

Network analysis revealing research focus of the German Congress of Orthopedics and Trauma Surgery 2021.

Technology and health care : official journal of the European Society for Engineering and Medicine·2026
Same author

Preemptive Antibiotic Administration in Open Fractures.

Zeitschrift fur Orthopadie und Unfallchirurgie·2026
Same author

Mandatory Labelling of CMR Substances in Implants and Their Impact on Orthopedics and Trauma Surgery.

Zeitschrift fur Orthopadie und Unfallchirurgie·2026
Same author

Sleep and Haemophilia-A Case-Control Analysis of Associated Factors.

Haemophilia : the official journal of the World Federation of Hemophilia·2026
Same author

AI-Generated Antibiotic Therapies for Acute Periprosthetic Joint Infections with Implant Retention in Comparison with an Interdisciplinary Team.

Antibiotics (Basel, Switzerland)·2026

Related Experiment Video

Updated: Jun 17, 2025

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
07:13

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities

Published on: October 27, 2023

1.1K

Evaluating multimodal AI in medical diagnostics.

Robert Kaczmarczyk1, Theresa Isabelle Wilhelm2, Ron Martin3

  • 1Department of Dermatology and Allergy, School of Medicine, Technical University of Munich, Munich, Germany.

NPJ Digital Medicine
|August 7, 2024
PubMed
Summary

This study assessed AI models against human intelligence for clinical diagnostic questions. While Claude 3 AI showed high accuracy, collective human decision-making remained superior, highlighting AI

More Related Videos

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K
Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
08:20

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images

Published on: October 27, 2023

1.4K

Related Experiment Videos

Last Updated: Jun 17, 2025

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
07:13

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities

Published on: October 27, 2023

1.1K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K
Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
08:20

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images

Published on: October 27, 2023

1.4K

Area of Science:

  • Artificial Intelligence
  • Medical Diagnostics
  • Clinical Decision Support

Background:

  • Multimodal AI models are increasingly explored for clinical applications.
  • Evaluating AI performance against human expertise is crucial for diagnostic tools.
  • The NEJM Image Challenge provides a benchmark for medical image interpretation.

Purpose of the Study:

  • To compare the accuracy and responsiveness of multimodal AI models with human collective intelligence.
  • To identify the strengths and limitations of current AI in clinical diagnostics.
  • To analyze AI performance variations based on question complexity and image characteristics.

Main Methods:

  • Utilized the NEJM Image Challenge dataset for evaluating AI models.
  • Assessed accuracy and responsiveness of various multimodal AI models (e.g., Claude 3, GPT-4 Vision Preview).
  • Compared AI performance against aggregated human expert decision-making.

Main Results:

  • Anthropic's Claude 3 family achieved the highest accuracy among tested AI models.
  • Claude 3 surpassed average human accuracy, but collective human decision-making outperformed all AI models.
  • GPT-4 Vision Preview demonstrated selective response patterns, favoring easier questions and smaller images.

Conclusions:

  • Multimodal AI shows significant potential in clinical diagnostics but has current limitations.
  • Human collective intelligence remains the benchmark for complex diagnostic tasks.
  • AI model performance is influenced by question difficulty and image properties, requiring further research.