Related Experiment Video
Updated: Jul 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams
1Hubert Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA USA.
Abstract:
LLMs have demonstrated remarkable ability in answering medical examinations. However, whether their performance remains stable under stress evaluations is unknown. We used PeruMedQA (n=8,380), a multiple-choice question-answering dataset, and ten medical LLMs. The stress test consisted of randomly shuffling the multiple-choice answers. Using paired t-tests and Wilcoxon tests, we compared LLM accuracy on the original versus the shuffled exams. MedGemma 27B, OctoMed-7B, and Meditron 7B did not exhibit statistically significant differences, overall and stratified by year/specialty. The largest non-significant differences for these LLMs were -2.82, -3.15, and -5.96 percentage points, respectively. These three LLMs may represent robust options for AI applications in Spanish-speaking Latin America.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s40670-026-02780-x.
