对葡萄牙语的开源大型语言模型进行比较Revalida多选题问题
João Victor Bruneti Severino1,2, Pedro Angelo Basei de Paula1, Matheus Nespolo Berger1
1Federal University of Parana, Curitiba, Brazil.
BMJ health & care informatics
|February 25, 2025
概括
这项研究评估了巴西医疗执照考试中的31个大型语言模型 (LLM). 像GPT-4o这样的顶级专有模型超过了人类的性能,而一些中型LLM也显示出了强的结果.
科学领域:
- 人工智能的人工智能
- 医学教育 医学教育
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 在各个领域都显得有前途.
- 评估LLM在医学等专业领域的表现至关重要.
- 巴西国家医学检查Revalida作为医学知识的严格基准.
研究的目的:
- 评估领先的大型语言模型 (LLM) 在葡萄牙语验证的医学知识测试中的表现.
- 在高风险的医学检查背景下,比较开源和专有LLM的功能.
主要方法:
- 对31个大型语言模型 (LLM) 的全面评估,包括23个开源和8个专有模型.
- 模型在巴西国家医学检查中进行了测试,包括399个多选择题.
- 通过Revalida基准的成功率来衡量业绩.
主要成果:
- 专有型号GPT-4o (86.8%) 和Claude Opus (83.8%) 的成功率最高.
- 在开源模型中,Llama 3 70B (77.5%),Mixtral 8×7B (63.7%) 和Llama 3 8B (53.9%) 显示出基于尺寸的不同性能.
- 十个LLM在Revalida基准上表现超过了人类水平.
结论:
- 几种大型语言模型 (LLM) 在Revalida医学检查中实现了专家级别的性能.
- 模型大小通常与性能相关,但一些中型LLM的表现优于较大的LLM.
- 在医学知识评估方面,LLM显示出显著的潜力,尽管对一些模型来说,连贯性仍然存在挑战.
更多相关视频
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.1K
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
385
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Multiple Comparison Tests
3.8K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.8K
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
