在提取认知考试日期和分数时评估大型语言模型.
Hao Zhang1, Neil Jethani1, Simon Jones1
1NYU Grossman School of Medicine, New York, New York, United States of America.
PLOS digital health
|December 11, 2024
概括
评估了ChatGPT和LlaMA-2以从临床笔记中提取认知测试分数. 聊天GPT在提取MMSE和CDR数据方面表现出卓越的准确性,对临床应用显示出更高的可靠性.
科学领域:
- 人工智能在医学中的应用
- 自然语言处理自然语言处理.
- 临床信息学 临床信息学
背景情况:
- 大型语言模型 (LLM) 的可靠性对于临床任务至关重要.
- 评估LLM用于提取特定的临床信息,例如认知测试结果,对于其安全实施至关重要.
研究的目的:
- 评估两个最先进的LLM,ChatGPT (GPT-4) 和 LlaMA-2在从电子健康记录中提取与认知测试 (MMSE,CDR,MoCA) 相关的临床信息方面的性能.
- 为了比较ChatGPT和Llama-2在识别和提取认知测试分数和日期方面的准确性,灵敏度和精度.
主要方法:
- 一个包含135,307条临床注释的数据集被策划,其中34,465条符合MMSE,CDR或MoCA提及的纳入标准.
- 765个笔记由ChatGPT和Llama-2处理,LLM输出经过专家审查.
- 性能指标包括准确度,灵敏度,精度和评价者之间的一致性 (Fleiss' Kappa) 按照TRIPOD指南进行计算.
主要成果:
- 在MMSE提取 (83%对Llama-2的66.4%) 和CDR提取 (87.1%对Llama-2的74.5%) 中,ChatGPT获得了更高的精度.
- 与Lama-2相比,ChatGPT在MMSE和CDR数据提取方面表现出明显更好的灵敏度和精度.
- 定性分析显示,与Llama-2相比,ChatGPT的幻觉和错误日期报告的情况较少.
结论:
- 在从临床笔记中提取认知测试信息方面,ChatGPT表现出优于 LlaMA-2 的性能,这表明它有临床使用的潜力.
- 像ChatGPT这样的LLM可以帮助痴呆症研究和治疗或临床试验的患者鉴定.
- 严格的验证对于了解医疗保健环境中的LLM能力和局限性至关重要.
更多相关视频
06:58Highlighting and Reducing the Impact of Negative Aging Stereotypes During Older Adults' Cognitive Testing
Published on: January 24, 2020
7.3K
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.1K
相关概念视频
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Language and Cognition
323
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
323
