大型语言模型在医疗保健应用中的输入变量下的性能:数据集开发和实验评估
Saubhagya Joshi1, Monjil Mehta2, Sarjak Maniar2
1Library and Information Sciences, School of Communication & Information, Rutgers University, 4 Huntington St, New Brunswick, NJ, 08901, United States, +1 (848) 932-7500.
JMIR AI
|February 20, 2026
概括
医疗保健中的大型语言模型 (LLM) 显示出令人惊的稳定性,可以应对常见的输入错误,如打字错误和同声语言. 然而,编辑会显著降低LLM的性能,强调在临床应用中需要谨慎设计.
科学领域:
- 医疗保健中的人工智能
- 自然语言处理自然语言处理.
- 临床信息学 临床信息学
背景情况:
- 大型语言模型 (LLM) 在医疗保健中越来越多地用于患者护理和决策.
- 缺乏完整临床数据的LLM的可靠性尚未得到充分理解.
- 数据缺陷在临床文档和患者生成的信息中很常见.
研究的目的:
- 研究输入干扰对健康应用中的LLM性能的影响.
- 比较不同类型和级别的干扰的影响.
- 分析与健康相关的和与健康无关的术语的差异性影响.
主要方法:
- 在3个与健康相关的任务中对3个LLM进行系统评估.
- 使用了一套具有人类变异的新型数据集:编辑,同音和排版错误.
- 在各种扰动级别下评估性能.
主要成果:
- 在LLM中,对常见的输入变化表现出了显著的稳定性;在55%以上的案例中,性能是稳定的或有所改善.
- 较低的扰动水平有时会导致性能提高 (14.07%).
- 删除比其他变化更不利于LLM的表现.
结论:
- 使用LLM的医疗保健应用程序必须考虑到输入变化和数据质量.
- 对不完美的输入的稳定性对于临床环境中的LLM可靠性至关重要.
- 调查结果为开发弹性AI工具和改善医疗保健中的LLM绩效提供了见解.
相关概念视频
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K
Variability: Analysis
547
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
547
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
277
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
277


