加强隐私保护可部署的大型语言模型用于外科手术期间并发症检测:一个有针对性的战略与LoRA微调
Shaowei Gao1, Xu Zhao2, Lihui Chen3
1Department of Anesthesiology, First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China. gaoshw5@mail.sysu.edu.cn.
NPJ digital medicine
|December 13, 2025
概括
这项研究表明,有针对性的提示工程和低级适应 (LoRA) 微调可以将小型开源语言模型转化为专家级别的工具,用于识别和分级外科手术后的并发症,克服手动检测和当前人工智能部署的局限性. 这些优化的模型实现了专家级准确性,在复杂的文档中保持性能,并允许在保留数据主权的同时进行本地部署,为医疗保健提供了实用的解决方案.
科学领域:
- 人工智能在医学中的应用
- 临床自然语言处理 临床自然语言处理
- 医疗保健信息学 医疗保健信息学
背景情况:
- 外科手术期间的并发症对全球健康构成重大挑战,手动检测方法显示大量报告不足 (27%) 和错误分类率.
- 部署临床大型语言模型 (LLM) 受到数据隐私问题,高计算成本和局部部署模型的低于最佳性能的阻碍.
研究的目的:
- 开发和验证一个框架,使用有针对性的提示工程和低级适应 (LoRA) 微调来增强较小的,开源的LLM的诊断能力,用于外科手术后并发症检测和严重程度分级.
- 评估优化的LLM与人类专家的性能,并评估其对临床文档复杂性变化的稳定性.
主要方法:
- 建立了一个双中心验证框架,以同时识别和分级22个不同的外科手术期间并发症的严重程度.
- 针对性的提示工程,包括思维链提示和LoRA微调,应用于较小的开源LLM.
- 使用F1分数在不同的文档长度四分位数中评估性能,并将AI模型和人类专家进行比较.
主要成果:
- 优化的LLM,特别是4B和8B参数模型,在识别和分类外科手术后并发症方面表现出专家级准确性,8B模型超过了人类专家的表现 (F1>0.70).
- 目标策略显著改善了模型性能 (4B模型的ΔF1=0.256),从LoRA获得了进一步的收益 (4B模型的ΔF1=0.103),在外部验证时将4B模型的微F1提高到0.64.
- 人工智能模型对文档复杂性表现出优越的稳定性,保持高F1分数 (F1>0.64),而人类专家的性能显著下降 (从0.73到0.45).
结论:
- 针对性提示工程与LoRA微调相结合,有效地将较小的开源LLM转化为高性能临床诊断工具,用于术后并发症.
- 这些优化的小型模型为资源有限的医疗保健环境提供了实用解决方案,使专家级准确度能够在当地部署和保留数据主权的情况下实现.
- 该研究强调了在临床环境中利用人工智能的可行途径,解决了复杂性检测和管理的关键挑战.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Types of Errors: Detection and Minimization
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
