挪威临床文本中的代码预测:模型开发和评估研究
Phuong Dinh Ngo1,2, Miguel Ángel Tejedor Hernández1,3, Taridzo Chomutare1,4
1Norwegian Centre for E-health Research, University Hospital of Northern Norway, P.O. Box 35, N-9038, Tromsø, Norway, 47 92699162.
JMIR AI
|August 25, 2025
概括
新的挪威BERT模型NorDeClin-BERT显著提高了国际疾病统计分类第十版 (ICD-10) 的编码精度. 对于挪威临床文本来说,特定领域的预培训提高了性能.
科学领域:
- 自然语言处理 (NLP)
- 机器学习
- 卫生信息学
背景情况:
- 准确的国际疾病统计分类第十次修订 (ICD-10) 编码对于医疗保健运作至关重要,但手动流程容易出现错误和效率低下.
- 目前用于ICD-10编码的NLP模型主要集中在英语,为挪威临床文本创造了研究缺口.
- 需要针对挪威医疗保健系统的自动化ICD-10编码解决方案.
研究的目的:
- 介绍NorDeClin-BERT,这是一个针对特定领域的挪威BERT模型,用于增强医学语言理解.
- 评估特定领域预训练和模型大小对ICD-10代码分类性能的影响.
- 将Nordeclin-BERT与挪威ICD-10编码的通用和跨语言BERT模型进行比较.
主要方法:
- 在ClinCode Gastro Corpus的880万个挪威临床笔记上预先训练了两个版本的NorDeClin-BERT (基本和大).
- 微调了ICD-10诊断码预测的模型.
- 与SweDeClin-BERT,ScandiBERT,NorBERT3-base和NorBERT3-large进行比较,使用准确度,精度,回忆和F1分数.
主要成果:
- 在分类ICD-10代码方面,NorDeClin-BERT版本的表现都超过了挪威的一般BERT模型和瑞典的临床BERT模型.
- 在所有评估指标中,NorDeClin-BERT-large取得了最高的表现,证明了特定领域预培训和模型能力的好处.
- 瑞典的临床模型显示出有限的可转移性,强调了对挪威的临床预训的需要.
结论:
- 在挪威胃肠病学中,NorDeClin-BERT具有显著的潜力,可以改善ICD-10代码分类,简化文档和减少行政负担.
- 这项研究确立了NorDeClin-BERT作为挪威医学NLP和ICD-10编码的最先进模型,为研究奠定了新的基础.
- 未来的研究应探索先进的领域适应,外部知识整合和跨医院通用性,以获得更广泛的临床应用.
相关概念视频
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


