Related Experiment Video
Updated: Aug 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions:
Moshe Carroll1, Sabrina Kentis2, Hannah Kareff2
1Albert Einstein College of Medicine, Department of Medicine, Montefiore Medical Center, United States, New York.
Background:
Large language models are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist.
Objectives:
To evaluate and compare the accuracy and response consistency of three leading large language models-ChatGPT-5, Gemini 2.5, and Grok 4-on Allergy/Immunology United States Medical Licensing Examination Step 1-style questions under different prompt conditions.
Methods:
Thirty-five United States Medical Licensing Examination Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy and variability was measured using Shannon entropy. Mixed effects models tested effects of model, prompt condition, and question difficulty.
Results:
Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability.
Conclusions:
On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.
