Related Experiment Video
Updated: May 25, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking open-source large language models on Portuguese Revalida multiple-choice questions.
João Victor Bruneti Severino1,2, Pedro Angelo Basei de Paula1, Matheus Nespolo Berger1
1Federal University of Parana, Curitiba, Brazil.
This study evaluated 31 large language models (LLMs) on the Brazilian medical licensing exam. Top proprietary models like GPT-4o surpassed human performance, while some medium-sized LLMs also showed strong results.
Area of Science:
- Artificial Intelligence
- Medical Education
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in various fields.
- Evaluating LLM performance in specialized domains like medicine is crucial.
- The Revalida, Brazil's national medical examination, serves as a rigorous benchmark for medical knowledge.
Purpose of the Study:
- To assess the performance of leading large language models (LLMs) on validated medical knowledge tests in Portuguese.
- To compare the capabilities of open-source and proprietary LLMs in a high-stakes medical examination context.
Main Methods:
- A comprehensive evaluation of 31 large language models (LLMs), including 23 open-source and 8 proprietary models.
- Models were tested on the national Brazilian medical examination, comprising 399 multiple-choice questions.
- Performance was measured by success rates on the Revalida benchmark.
Main Results:
- Proprietary models GPT-4o (86.8%) and Claude Opus (83.8%) achieved the highest success rates.
- Among open-source models, Llama 3 70B (77.5%), Mixtral 8×7B (63.7%), and Llama 3 8B (53.9%) showed varying performance based on size.
- Ten LLMs demonstrated performance exceeding the human level on the Revalida benchmark.
Conclusions:
- Several large language models (LLMs) achieved expert-level performance on the Revalida medical examination.
- Model size generally correlated with performance, but some medium-sized LLMs outperformed larger ones.
- LLMs show significant potential in medical knowledge assessment, though challenges with coherence remain for some models.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Reliability and Validity
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Quantifying and Rejecting Outliers: The Grubbs Test
Improving Translational Accuracy
Expected Frequencies in Goodness-of-Fit Tests
Goodness-of-Fit Test