与基础医学科学考试中的大语言模型准确性相关的因素:横截面研究
Naritsaret Kaewboonlert1, Jiraphon Poontananggul1, Natthipong Pongsuwan1
1Institute of Medicine, Suranaree University of Technology, 111 University Avenue, Nakhon Ratchasima, 30000, Thailand, 66 44223956.
JMIR medical education
|January 23, 2025
概括
与其他大型语言模型 (LLM) 相比,GPT-4和微软Bing在基础医学科学考试中表现出更高的准确性. 模型性能与问题难度相关,更容易的问题为先进的人工智能工具带来更高的准确性.
科学领域:
- 医疗教育中的人工智能
- 在医疗保健中的自然语言处理.
- 科学知识的计算语言学
背景情况:
- 人工智能 (AI) 越来越多地被用于医学教育.
- 大型语言模型 (LLM) 在医学评估中的准确性是积极研究的领域.
- 内容验证和模型性能依赖于培训数据和优化.
研究的目的:
- 评估领先的LLM (GPT-3.5,GPT-4,Google Bard,微软Bing) 在回答基础医学科学考试问题时的准确性.
- 确定影响该领域LLM准确性的因素.
主要方法:
- 使用了与泰国第一步国家医疗执照考试一致的多选题.
- 收集关于问题的难度,歧视指数和其他特征的数据.
- 在GPT-3.5,GPT-4,Microsoft Bing和Google Bard中输入问题,通过多变量逻辑回归分析准确度.
主要成果:
- GPT-4获得了最高的准确度 (89.07%),其次是微软的Bing (83.69%),GPT-3.5 (67.02%) 和谷歌的Bard (63.83%).
- 在问题难度和LLM绩效之间观察到显著的相关性,更容易的问题得到了更准确的答案.
- 对大多数LLM来说,模型准确性与问题长度,负面措辞或临床场景的相关性有限.
结论:
- 与GPT-3.5和谷歌Bard相比,GPT-4和微软Bing在基础医学科学评估中表现出更高的准确性.
- 法学士准确度受到问题难度的重大影响,在更容易的项目上表现更好.
- 准确的LLM如GPT-4和Bing可以作为医学科学的宝贵教育工具.
关键词:
在这里,我们可以看到AIAIAI.聊天GPT 聊天GPT 在线聊天谷歌谷歌谷歌谷歌是什么意思在法学士 (LLM) 课程中.准确度 准确度 准确度 准确度 准确度人工智能的人工智能是人工智能.评估评估的评估评估的评估.基础医学科学考试 考试 基础医学科学考试跨截面研究是跨截面研究.数据集数据集数据集.难度指数难度指数大型语言模型医学教育 医学教育医学科学 医学科学 医学科学业绩表现表现的表现表现是什么工具 工具 工具 工具更多相关视频
相关概念视频
Language and Cognition
320
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
320
Improving Translational Accuracy
2.5K
2.5K
Sensitivity, Specificity, and Predicted Value
177
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
177
Language Development
314
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
314


