使用共识大语言模型方法提高结构化数据提取的可靠性和准确性 - - 多发性硬化症的用例描述.
Philip Lennart Poser1, Rafael Klimas1, Justus Luerweg1
1Department of Neurology, St. Josef-Hospital, Ruhr-University Bochum, Bochum, Germany.
Frontiers in artificial intelligence
|March 2, 2026
概括
大型语言模型 (LLM) 现在可以分析用于多发性硬化症 (MS) 研究的临床报告. 使用LLM的共识方法实现了与人类专家可比的准确性,从而实现了高效的数据分析.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 临床数据管理 临床数据管理
背景情况:
- 缺乏临床数据文档的标准化,阻碍了大规模的回顾性研究.
- 手动提取数据是耗时的,昂贵的,容易产生偏见.
- 使用大语言模型 (LLM) 开发了一种半自动化方法,以应对多发性硬化症 (MS) 门诊报告中的这些挑战.
研究的目的:
- 开发和评估一种半自动化方法,使用LLMs从MS门诊报告中提取结构化数据.
- 为了比较LLM生成数据的准确性与神经病学专家的手动评估.
- 评估LLM共识方法在提高数据提取质量的有效性.
主要方法:
- 利用商业的LLM (OpenAI,Anthropic,Google) 在30个匿名的MS门诊报告上进行零射击学习.
- 通过结合来自三个不同模型的输出,实施了LLM共识机制.
- 在几次运行中代精制提示,并根据参考标准评估错误率.
- 计算了LLM共识和神经学家输出的真错率,仅考虑内容偏差.
主要成果:
- 快速的工程代导致了LLM错误率的显著降低.
- 该LLM共识方法克服了与个别LLM观察到的绩效上限.
- 该LLM共识实现了1.48%的真错率,与神经病学专家 (大约. 2%). 这是一个很好的方法.
结论:
- 开发的基于LLM的方法提供了一种快速,可靠和可访问的方式来分析大量非结构化临床数据.
- 通过LLM共识显著提高了输出质量,使其与专家手动数据创建相提并论.
- 虽然对时间和成本效益有希望,但在科学研究中严格验证基于LLM的方法仍然至关重要.
更多相关视频
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
11.2K
05:44Author Spotlight: Creating a Versatile Experimental Autoimmune Encephalomyelitis Model Relevant for Both Male and Female Mice
Published on: October 13, 2023
2.5K
相关概念视频
Improving Translational Accuracy
3.7K
3.7K
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
