Related Experiment Video
Updated: May 31, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models on the Brazilian National Medical Education Examination: Comparative Benchmark
Francys de Luca Fernandes da Silva1, Eduardo Augusto Roeder2, João Victor Bruneti Severino1,2
1R. Imac. Conceição, 1155 - Prado Velho, Pontifícia Universidade Católica do Paraná, Curitiba, Paraná, Brazil.
A specialized Brazilian Portuguese large language model (LLM) outperformed commercial LLMs on the 2026 Brazilian National Medical Education Examination (ENAMED 2026). This study provides crucial insights into LLM performance in high-stakes medical assessments.
Area of Science:
- Artificial Intelligence in Medical Education
- Natural Language Processing for Healthcare
- Large Language Model (LLM) Benchmarking
Background:
- Most large language model (LLM) performance data for medical education comes from English materials.
- The effectiveness of frontier commercial LLMs versus Brazilian Portuguese domain-specialized systems on Brazilian medical exams is unknown.
Purpose of the Study:
- To compare the performance of 9 frontier commercial LLMs and 1 Brazilian Portuguese domain-specialized system (Charcot) on the 2026 Brazilian National Medical Education Examination (ENAMED 2026).
- To identify systematic errors between models as a quality indicator.
Main Methods:
- 100 items from ENAMED 2026 were administered to 10 LLMs using identical Portuguese prompts.
- Primary outcome was mean accuracy; secondary outcomes included convergence error (CE), normalized mean response time (NMRT), and intermodel agreement.
- Statistical analyses included Shapiro-Wilk, Levene, Kruskal-Wallis, Dunn-Holm, and a binomial generalized linear mixed model.
Main Results:
- The Brazilian Portuguese system (Charcot) achieved the highest accuracy (96.97%), outperforming commercial LLMs.
- Accuracy varied significantly between models, with Charcot being statistically indistinguishable from the top commercial cluster.
- High intermodel agreement was observed (Fleiss κ=0.852), and CE prospectively identified a rectified exam item.
Conclusions:
- A Brazilian Portuguese domain-specialized LLM ranked highest on ENAMED 2026, performing comparably to frontier commercial models.
- Findings offer black-box performance evidence, not mechanistic proof of specialization.
- Convergence error (CE) shows promise as a quality assurance tool for LLM performance in medical examinations.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy