Related Experiment Video
Updated: Jul 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams
1Hubert Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA USA.
Medical Science Educator
|July 13, 2026
Summary
Ten medical large language models (LLMs) were tested on the PeruMedQA dataset under a stress test. MedGemma 27B, OctoMed-7B, and Meditron 7B showed stable performance, indicating their potential for AI applications in Latin America.
Area of Science:
- Medical Informatics
- Artificial Intelligence
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in medical question answering.
- LLM performance under stress conditions, such as answer shuffling, is not well-understood.
- Evaluating LLM robustness is crucial for reliable AI applications in healthcare.
Purpose of the Study:
- To assess the stability of ten medical LLMs' performance on a medical question-answering dataset under a stress test.
- To identify LLMs that maintain accuracy when multiple-choice answers are randomly shuffled.
- To determine the suitability of specific LLMs for AI applications in Spanish-speaking regions.
Main Methods:
- Utilized the PeruMedQA dataset (n=8,380), a multiple-choice medical question-answering dataset.
- Implemented a stress test involving random shuffling of multiple-choice answers for each question.
- Compared LLM accuracy on original versus shuffled exams using paired t-tests and Wilcoxon tests.
Main Results:
- MedGemma 27B, OctoMed-7B, and Meditron 7B demonstrated no statistically significant performance differences under the stress test.
- The largest non-significant accuracy differences for these three LLMs were -2.82%, -3.15%, and -5.96%, respectively.
- Performance stability was observed overall and when stratified by medical year and specialty.
Conclusions:
- MedGemma 27B, OctoMed-7B, and Meditron 7B exhibit robust performance, suggesting their reliability in medical AI applications.
- These LLMs are potential candidates for deployment in Spanish-speaking Latin American healthcare settings.
- Further research can explore additional stress factors to comprehensively evaluate LLM stability.
