PeruMedQA: A Stress Evaluation Using Ten Large Language Models to Answer Medical Exams

Rodrigo M Carrillo-Larco1,2

  • 1Hubert Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA USA.

Summary

Ten medical large language models (LLMs) were tested on the PeruMedQA dataset under a stress test. MedGemma 27B, OctoMed-7B, and Meditron 7B showed stable performance, indicating their potential for AI applications in Latin America.

Related Concept Videos