在提取认知考试日期和分数时评估大型语言模型
Hao Zhang1, Neil Jethani1, Simon Jones1
1NYU Grossman School of Medicine.
medRxiv : the preprint server for health sciences
|February 26, 2024
概括
评估了ChatGPT和LlaMA-2以从临床笔记中提取认知测试分数. 与LlaMA-2相比,ChatGPT在识别MMSE和CDR得分和日期方面表现出更高的准确性.
科学领域:
- 医疗保健中的人工智能
- 自然语言处理用于临床数据提取.
- 医疗信息学 医疗信息学
背景情况:
- 大型语言模型 (LLM) 在医疗应用中越来越重要.
- 确保医疗保健中LLM的可靠性至关重要,以防止错误的临床信息.
- 这项研究的重点是提取认知测试数据,特别是迷你精神状态检查 (MMSE) 和临床痴呆评分 (CDR) 的得分.
研究的目的:
- 评估两个先进的LLM,ChatGPT (GPT-4) 和 LlaMA-2在提取MMSE和CDR分数方面的表现.
- 评估这些LLM在从临床笔记中提取与认知测试得分相关的日期的准确性.
- 为了比较ChatGPT与Lama-2在临床信息提取任务中的有效性.
主要方法:
- 在2010年1月12日至2023年5月24日期间,分析了34465份提到MMSE,CDR或蒙特利尔认知评估 (MoCA) 的临床笔记.
- 765条笔记由ChatGPT和Llama-2处理,结果由22名专家审查.
- 绩效指标包括准确性,精度,回忆力,真假负率和评价者之间的协议 (Fleiss' Kappa),遵守TRIPOD指南.
主要成果:
- 与Lama-2 (66.4%) 相比,ChatGPT在MMSE提取方面获得了更高的精度 (83%),具有更高的灵敏度和精度.
- 对于CDR提取,ChatGPT也表现优于Lama-2,显示出更高的精度 (87.1%与74.5%) 和灵敏度 (84.3%与39.7%).
- 定性分析显示,与Llama-2相比,ChatGPT的幻觉和报告错误较少.
结论:
- 在从临床笔记中提取认知测试分数和日期时,ChatGPT的准确性和可靠性明显高于Lama-2.
- 这些发现表明,像ChatGPT这样的LLM可以在痴呆症研究和临床护理中成为有价值的工具,有助于确定患者进行试验和治疗.
- 对LLM能力和局限性的持续严格评估对于其安全有效地融入医疗实践至关重要.
更多相关视频
06:58Highlighting and Reducing the Impact of Negative Aging Stereotypes During Older Adults' Cognitive Testing
Published on: January 24, 2020
7.3K
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.2K
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Language and Cognition
346
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
346
