麻醉学的人工智能委员会式考试问题:大型语言模型的作用
Adnan A Khan1, Rayaan Yunus1, Mahad Sohail1
1Department of Anesthesia, Critical Care, and Pain Medicine, Beth Israel Deaconess Medical Center, Harvard Medical School Boston, MA.
像大型语言模型 (LLM) 这样的人工智能工具在麻醉学教育中显示出潜力,但目前在董事会考试中失败. GPT-4的表现最好,但没有一个获得通过分数,表明需要进一步发展.
科学领域:
- 医学教育 医学教育
- 人工智能的人工智能
- 麻醉学 麻醉学
背景情况:
- 大型语言模型 (LLM) 是新兴的人工智能工具,具有潜在的医疗应用.
- 在麻醉学教育中LLMs的可靠性仍然未被探索.
- 评估LLM绩效对于理解它们在医学培训中的有用性至关重要.
研究的目的:
- 评估公开可用的麻醉学教育LLM的表现.
- 评估ChatGPT (GPT-3.5,GPT-4) 和谷歌Bard在麻醉学会考试问题上的可靠性.
主要方法:
- 进行了探索性前性审查.
- 三个LLM (GPT-3.5,GPT-4,Bard) 在麻醉委员会审查书中的884个问题上进行了测试.
主要成果:
- 总体而言,正确答案率为47.9% (GPT-3.5),69.4% (GPT-4) 和45.2% (巴德).
- GPT-4的表现明显优于GPT-3.5和Bard (p < 0.001) 的表现.
- 没有LLM达到董事会认证的70%通过门.
结论:
- 目前的LLM没有足够先进的麻醉学委员会考试成功.
- 虽然GPT-4表现出卓越的性能,但缺乏一致的,医学上合理的解释.
- 未来的特定领域培训可能会提高医学教育中的LLM实用性.
更多相关视频
13:12Translational Brain Mapping at the University of Rochester Medical Center: Preserving the Mind Through Personalized Brain Mapping
Published on: August 12, 2019
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
相关概念视频
Stages of General Anesthesia
Local Anesthetics: Clinical Application as Spinal Anesthesia
General Anesthesia: Overview
General anesthesia induces unconsciousness in the whole body, while the others target specific areas or sensations. It is administered to minimize adverse effects, maintain...
Local Anesthetics: Clinical Application as Surface, Infiltration, and Conduction Block Anesthesia
Inhalational Anesthetics: Overview
Skeletal Muscle Relaxants: Therapeutic Uses
