Related Experiment Video
Updated: Jun 4, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
500
Large Language Models in Worldwide Medical Exams: Platform Development and Comprehensive Analysis
Hui Zong1, Rongrong Wu1, Jiaxue Cha2
1Joint Laboratory of Artificial Intelligence for Critical Care Medicine, Department of Critical Care Medicine and Institutes for Systems Genetics, Frontiers Science Center for Disease-related Molecular Network, West China Hospital, Sichuan University, Chengdu, China.
Journal of Medical Internet Research
|December 27, 2024
Summary
MedExamLLM evaluates large language models (LLMs) on global medical exams, revealing GPT-4
Area of Science:
- Artificial Intelligence in Medical Education
- Natural Language Processing
- Health Professions Education
Background:
- Large language models (LLMs) show promise for medical education and assessment.
- Global performance of LLMs across diverse medical examinations is not well understood.
Purpose of the Study:
- Introduce MedExamLLM, a platform to systematically evaluate LLM performance on worldwide medical exams.
- Compile and analyze LLM performance data across regions, languages, and contexts.
- Provide a resource for advancing AI integration in medical education.
Main Methods:
- Systematic PubMed search (April 25, 2024) for English-language, peer-reviewed studies evaluating LLMs on medical exams.
- Independent screening by two researchers, data curation on exam details, LLM performance, and references.
- Integration of curated data into the MedExamLLM platform for visualization and analysis.
Main Results:
- 193 articles analyzed, covering 16 LLMs and 198 medical exams in 28 countries and 15 languages (2009-2023).
- Generative Pretrained Transformer (GPT) models, particularly GPT-4, showed superior performance.
- Significant variability in LLM capabilities observed across geographic and linguistic contexts.
Conclusions:
- MedExamLLM is an open-source platform for evaluating LLM performance on global medical exams.
- It serves as a resource for educators, researchers, and developers integrating AI in medicine.
- Limitations include potential data bias and exclusion of non-English literature; future research is needed.
Keywords:
AIChatGPTLLMsartifical intelligencegenerative pretrained transformerlarge language modelsmedical educationmedical exam
