Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jul 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams.

Rodrigo M Carrillo-Larco1,2

  • 1Hubert Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA USA.

Medical Science Educator
|July 13, 2026
PubMed
Summary

Ten medical large language models (LLMs) were tested on the PeruMedQA dataset under a stress test. MedGemma 27B, OctoMed-7B, and Meditron 7B showed stable performance, indicating their potential for AI applications in Latin America.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Epidemiology and risk factors of stroke in the Americas: a comprehensive narrative literature review.

Lancet regional health. Americas·2026
Same author

Implementation of stroke prevention: a review of challenges and opportunities in the Americas.

Lancet regional health. Americas·2026
Same author

Does Domain-Specific Retrieval Augmented Generation Help LLMs Answer Consumer Health Questions?

Proceedings of machine learning research·2026
Same author

Association of life-course stressful life events with later-life intrinsic capacity: A multi-cohort study.

The journal of nutrition, health & aging·2026
Same author

Age at type 2 diabetes diagnosis, mortality, and health loss in South Asians.

Diabetes research and clinical practice·2026
Same author

Role of retinal biomarkers in diabetes detection and risk prediction: A systematic scoping review.

Diabetes & metabolic syndrome·2026

Area of Science:

  • Medical Informatics
  • Artificial Intelligence
  • Natural Language Processing

Background:

  • Large language models (LLMs) show promise in medical question answering.
  • LLM performance under stress conditions, such as answer shuffling, is not well-understood.
  • Evaluating LLM robustness is crucial for reliable AI applications in healthcare.

Purpose of the Study:

  • To assess the stability of ten medical LLMs' performance on a medical question-answering dataset under a stress test.
  • To identify LLMs that maintain accuracy when multiple-choice answers are randomly shuffled.
  • To determine the suitability of specific LLMs for AI applications in Spanish-speaking regions.

Main Methods:

  • Utilized the PeruMedQA dataset (n=8,380), a multiple-choice medical question-answering dataset.

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Related Experiment Videos

Last Updated: Jul 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

  • Implemented a stress test involving random shuffling of multiple-choice answers for each question.
  • Compared LLM accuracy on original versus shuffled exams using paired t-tests and Wilcoxon tests.
  • Main Results:

    • MedGemma 27B, OctoMed-7B, and Meditron 7B demonstrated no statistically significant performance differences under the stress test.
    • The largest non-significant accuracy differences for these three LLMs were -2.82%, -3.15%, and -5.96%, respectively.
    • Performance stability was observed overall and when stratified by medical year and specialty.

    Conclusions:

    • MedGemma 27B, OctoMed-7B, and Meditron 7B exhibit robust performance, suggesting their reliability in medical AI applications.
    • These LLMs are potential candidates for deployment in Spanish-speaking Latin American healthcare settings.
    • Further research can explore additional stress factors to comprehensively evaluate LLM stability.