Related Experiment Video
Updated: Jul 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of ChatGPT and Bard on the official part 1 FRCOphth practice questions
Thomas Fowler1, Simon Pullen2, Liam Birkett3
1Department of Medicine, Barking Havering and Redbridge University Hospitals NHS Trust, London, UK thomas.fowler6@nhs.net.
Background:
Chat Generative Pre-trained Transformer (ChatGPT), a large language model by OpenAI, and Bard, Google's artificial intelligence (AI) chatbot, have been evaluated in various contexts. This study aims to assess these models' proficiency in the part 1 Fellowship of the Royal College of Ophthalmologists (FRCOphth) Multiple Choice Question (MCQ) examination, highlighting their potential in medical education.
Methods:
Both models were tested on a sample question bank for the part 1 FRCOphth MCQ exam. Their performances were compared with historical human performance on the exam, focusing on the ability to comprehend, retain and apply information related to ophthalmology. We also tested it on the book 'MCQs for FRCOpth part 1', and assessed its performance across subjects.
Results:
ChatGPT demonstrated a strong performance, surpassing historical human pass marks and examination performance, while Bard underperformed. The comparison indicates the potential of certain AI models to match, and even exceed, human standards in such tasks.
Conclusion:
The results demonstrate the potential of AI models, such as ChatGPT, in processing and applying medical knowledge at a postgraduate level. However, performance varied among different models, highlighting the importance of appropriate AI selection. The study underlines the potential for AI applications in medical education and the necessity for further investigation into their strengths and limitations.
More Related Videos
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Comparing Experimental Results: Student's t-Test
Prismatic Beams: Problem Solving
The design begins with analyzing the beam as a free body to identify moments and force balances, thereby determining support reactions. Next, the...
Comparison between RL and RC circuits
Machines: Problem Solving II
Comparison Between Electrical And Gravitational Forces
Since both are inverse square law forces, the distance gets canceled when the ratio of the two forces is considered. Instead, the ratio of the electrical and gravitational forces depends on...

