Related Experiment Video
Updated: Jun 4, 2025

Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
Published on: December 8, 2023
Evaluating Bard Gemini Pro and GPT-4 Vision Against Student Performance in Medical Visual Question Answering:
Jonas Roos1, Ron Martin2, Robert Kaczmarczyk3
1Department of Orthopedics and Trauma Surgery, University Hospital of Bonn, Venusberg-Campus 1, 53127, Bonn, Germany, 49 228-287-14170.
Large language models (LLMs) show promise in medical image diagnostics, with GPT-4 Vision outperforming Bard Gemini. However, neither AI model matched student performance, highlighting areas for improvement in medical AI development.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Diagnostic Imaging Analysis
Background:
- Large language models (LLMs) are increasingly impacting medical research and education.
- LLMs with image recognition capabilities show potential in areas like radiological interpretation and exam assistance.
- Recent advancements have enhanced LLMs with crucial image analysis functionalities.
Purpose of the Study:
- To critically evaluate the effectiveness of LLMs in medical diagnostics and training.
- To assess the accuracy and utility of image-recognition-enhanced LLMs in answering medical licensing examination questions.
- To compare the performance of different LLMs against student benchmarks.
Main Methods:
- Analysis of 1070 image-based multiple-choice questions from the AMBOSS learning platform (605 English, 465 German).
- Utilized customized prompts in English and German to direct LLMs (GPT-4 1106 Vision Preview, Bard Gemini Pro) to interpret medical images and provide diagnoses.
- Compared LLM performance against student performance data (student passed mean, majority vote) using statistical analysis in Python.
Main Results:
- GPT-4 1106 Vision Preview achieved 56.9% accuracy, outperforming Bard Gemini Pro at 44.6% (P<.001).
- GPT-4 1106 Vision Preview had a higher unanswered rate (16.1%) than Bard Gemini Pro (4.1%).
- Both models performed better in German than English; however, student majority vote (94.5%) significantly outperformed both AI models.
Conclusions:
- LLMs like GPT-4 Vision and Bard Gemini demonstrate potential for medical visual question-answering and student support.
- Performance varies by language, with a noted advantage for German, and limitations exist for non-English content.
- Current LLM accuracy rates, especially compared to student responses, underscore the need for further optimization and understanding of their limitations in medical education.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025