评估大型语言模型在回答喘多选择和客观结构化临床检查问题方面的准确性
Pei Ye Li1, Andrea Gershon2, Andrew Kouri3
1Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, Toronto, Ontario, M5S 3K3, Canada.
大型语言模型 (LLM) 在回答成人喘问题方面显示出高准确度,ChatGPT模型表现异常出色. 这些先进的人工智能工具有望提高患者和临床医生的喘教育.
科学领域:
- 人工智能在医学中的应用
- 临床决策支持系统 临床决策支持系统
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 在临床环境中显示出潜力,但它们对喘知识的具体应用需要进一步调查.
- 最近的LLM进步 (例如,ChatGPT-5,Claude-3.7) 尚未对喘相关任务进行评估.
研究的目的:
- 评估各种大型语言模型 (LLM) 在回答成人喘多选择题 (MCQ) 和客观结构化临床检查 (OSCE) 中的准确性.
主要方法:
- 在5次代中,14个LLM聊天机器人被评估,使用116个成人喘MCQ和3个喘OSCE.
- 采用了通用线性混合效应模型来比较LLM准确性,考虑了问题类型 (MCQ与OSCE) 和模型特征 (通用与药物特定,开源与专有).
主要成果:
- 大多数LLM都取得了很好的MCQ准确度 (>85%),其中一些超过95%.
- 欧安组织的准确性在70-86%之间,ChatGPT-5排名最高.
- 在MCQ上,LLM的表现显著好于OSCE (71.1%) (92.8%),在面向患者的MCQ上 (96.3%) 和面向临床医生的MCQ (85.6%).
结论:
- 目前的LLM在喘相关的临床和以患者为中心的调查中显示出强大的准确性,特别是ChatGPT模型.
- 这些LLM对患者和临床医生来说都是未来喘教育计划的潜在有价值的工具.
更多相关视频
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
相关概念视频
Asthma-IV: Diagnostic and Management
Clinical Assessment for Asthma:
This is the first step in diagnosing and managing asthma. It includes:
Asthma-II: Pathophysiology and Classification
Additionally, environmental and genetic factors play crucial roles in determining an individual's susceptibility to asthma and the severity of their condition.
Critical processes in asthma pathophysiology include:
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Asthma-III: Symptoms and Complications
Classification of Asthma
Asthma-IV: Nursing Management
First, in...
