Related Experiment Video
Updated: Jun 28, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
550
Google Gemini and Bard artificial intelligence chatbot performance in ophthalmology knowledge assessment.
Andrew Mihalache1, Justin Grad2, Nikhil S Patil2
1Temerty Faculty of Medicine, University of Toronto, Toronto, ON, Canada.
Eye (London, England)
|April 13, 2024
Summary
Google Gemini and Bard show acceptable accuracy on ophthalmology board exam questions. Performance varied slightly by country, and chatbots sometimes confidently provided incorrect answers.
Area of Science:
- Artificial Intelligence in Medicine
- Ophthalmology Knowledge Assessment
Background:
- The rise of AI chatbots like ChatGPT necessitates evaluating their medical applications.
- Understanding the capabilities of Google Gemini and Bard in specialized medical fields is crucial.
Purpose of the Study:
- To assess the ophthalmology knowledge of Google Gemini and Bard.
- To evaluate their performance on board certification practice questions.
Main Methods:
- Google Gemini and Bard were tested on the EyeQuiz platform (150 multiple-choice questions).
- Evaluated metrics included accuracy, response length, time, and explanation quality.
- Secondary analysis included country-specific versions (US, Vietnam, Brazil, Netherlands).
Main Results:
- Overall accuracy for both chatbots was 71%.
- Country-specific versions showed minor accuracy variations (e.g., Vietnam Gemini: 74%, US Gemini: 71%).
- Differences in performance across countries were not statistically significant.
Conclusions:
- Google Gemini and Bard demonstrate acceptable performance in ophthalmology knowledge recall.
- Subtle international performance variations exist, though not statistically significant.
- Chatbots may offer confident explanations even for incorrect answers.

