Related Experiment Video
Updated: Jun 17, 2025

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
Evaluating multimodal AI in medical diagnostics
Robert Kaczmarczyk1, Theresa Isabelle Wilhelm2, Ron Martin3
1Department of Dermatology and Allergy, School of Medicine, Technical University of Munich, Munich, Germany.
Abstract:
This study evaluates multimodal AI models' accuracy and responsiveness in answering NEJM Image Challenge questions, juxtaposed with human collective intelligence, underscoring AI's potential and current limitations in clinical diagnostics. Anthropic's Claude 3 family demonstrated the highest accuracy among the evaluated AI models, surpassing the average human accuracy, while collective human decision-making outperformed all AI models. GPT-4 Vision Preview exhibited selectivity, responding more to easier questions with smaller images and longer questions.

