评估烟草使用障碍的大型语言模型的临床能力:多领域专家评估评估
Thiago P Fernandes1, Linnea Dahlgren2, Natanael A Santos1,3
1Department of Psychology, Perception, Neuroscience, and Behaviour Lab, Federal University of Paraiba, Joao Pessoa.
领先的大型语言模型 (LLM) 在戒烟支持方面显示出前景. GPT-4.5和Claude 3.5 Sonnet证明了临床能力,但临床医师的监督对所有AI干预至关重要.
科学领域:
- 医疗保健中的人工智能
- 数字健康干预措施 数字健康干预措施
- 临床决策支持系统 临床决策支持系统
背景情况:
- 烟草使用障碍 (TUD) 是全球主要的可预防的死亡原因.
- 由于劳动力短缺,TUD受限于符合指导方针的护理.
- 人工智能 (AI) 和大型语言模型 (LLM) 为扩大戒烟支持提供了潜在的解决方案.
研究的目的:
- 系统地评估五个领先的LLMs的临床准确性,安全性,准则遵守性和沟通质量.
- 在标准化戒烟情景中评估LLM绩效.
- 描述TUD治疗新兴AI系统的临床能力.
主要方法:
- 开发了84个临床细节,涵盖查,诊断,药物治疗,咨询和减少危害.
- 五个LLM的评估 (GPT-4.5,克劳德 3.5 索内特,双子座 2.5 专业,拉玛 3.1-70B,DeepSeek-V3) 由成医学专家.
- 评估标准包括临床准确性,遵循指南,安全性和临床实用性.
主要成果:
- GPT-4.5和克劳德3.5索内特获得了最高的综合分数,74-78%的响应被评为≥4.0.0.
- 这些顶级型号展示了卓越的安全性能 (88%的评级≥4.0).
- 双子 2.5 Pro 显示中等性能 (52% ≥4.0),而开放式重量模型 (Llama 3.1-70B,DeepSeek-V3) 落后.
结论:
- 所有评估的人工智能系统都表现出在戒烟辅导方面的能力.
- GPT-4.5和克劳德3.5索内特达到了适合监督临床使用的性能水平.
- 在以药物为基础的干预中,临床人员的监督是必不可少的;开放式权重模型需要进一步验证.
更多相关视频
05:12Chronic Intermittent Ethanol Vapor Exposure Paired with Two-Bottle Choice to Model Alcohol Use Disorder
Published on: June 23, 2023
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
相关概念视频
Drug Dependence
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Diagnostic and Statistical Manual of Mental Disorders (DSM)
CNS Depressants: Alcohol and Nicotine
