Related Experiment Video
Updated: May 15, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Performance of Anthropic Claude and OpenAI GPT Models in Basic Radiological Imaging Tasks
Cindy Nguyen1, Daniel Carrion2, Mohamed K Badawy1,2
1Department of Medical Imaging and Radiation Sciences, Monash University, Clayton, Victoria, Australia.
Publicly available artificial intelligence (AI) Vision Language Models (VLMs) show promise but are not yet ready for clinical radiology. Further development and standardized testing are needed for reliable performance in image interpretation.
Area of Science:
- Artificial Intelligence
- Radiology
- Medical Imaging
Background:
- Rapid advancements in AI Vision Language Models (VLMs) present opportunities to enhance radiology workflows.
- Evaluating the performance of these VLMs in radiological image interpretation is crucial for potential clinical integration.
Purpose of the Study:
- To assess the accuracy and consistency of publicly available VLMs, specifically Anthropic's Claude and OpenAI's GPT, in basic radiological image interpretation tasks.
- To evaluate model performance across multiple iterations to understand reliability.
Main Methods:
- Utilized subsets from ROCOv2 and MURAv1.1 datasets to test 6 VLMs.
- Inputted system prompts and images into each model thrice, comparing outputs to dataset captions.
- Analyzed accuracy in recognizing modality, anatomy, and detecting fractures, alongside output consistency.
Main Results:
- High accuracy (up to 100%) in modality recognition on ROCOv2.
- Anatomical recognition accuracy ranged from 61% to 85% across models.
- Claude-3.5-Sonnet demonstrated the highest anatomical recognition (57%) and consistency (83% anatomy, 92% fractures) on MURAv1.1; GPT-4o showed the best fracture detection (62%).
Conclusions:
- Current AI VLMs like Claude and GPT lack the necessary accuracy and reliability for clinical radiology integration.
- Ongoing research and development are essential, alongside standardized testing methodologies, to ensure dependable performance of AI in medical imaging.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
06:53Management of Respiratory Motion Artefacts in 18F-fluorodeoxyglucose Positron Emission Tomography using an Amplitude-Based Optimal Respiratory Gating Algorithm
Published on: July 23, 2020
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...