NorDeClinのドメイン固有の研修 - 病気の国際統計分類第10版のトランスフォーマーからの双方向エンコーダー表現 - ノルウェー臨床文献におけるコード予測: モデル開発と評価研究
Phuong Dinh Ngo1,2, Miguel Ángel Tejedor Hernández1,3, Taridzo Chomutare1,4
1Norwegian Centre for E-health Research, University Hospital of Northern Norway, P.O. Box 35, N-9038, Tromsø, Norway, 47 92699162.
JMIR AI
|August 25, 2025
まとめ
ノルウェーの新しいBERTモデルであるNorDeClin-BERTは,第10版国際統計疾患分類 (ICD-10) のコードの正確性を大幅に改善しています. ノルウェー語臨床テキストの一般的なモデルよりも,領域特有の予習がパフォーマンスを向上させます.
科学分野:
- 自然言語処理 (NLP)
- 機械学習
- 医療情報学
背景:
- 正確な国際統計疾病分類 第10版 (ICD-10) のコーディングは医療業務に不可欠ですが,手作業のプロセスは誤りになりやすく,非効率です.
- ICD-10のコード化のための既存のNLPモデルは主に英語に焦点を当てており,ノルウェー語の臨床テキストの研究のギャップを生み出しています.
- ノルウェー保健医療システムに合わせた自動化されたICD-10のコード化ソリューションが必要である.
研究 の 目的:
- NorDeClin-BERTを導入し,医療言語の理解を向上させるための分野特有のノルウェー語BERTモデルを導入する.
- ICD-10コード分類の性能に対する領域特有の予備訓練とモデルの大きさの影響を評価する.
- NorDeClin-BERTとノルウェーのICD-10コードの汎用および言語間のBERTモデルを比較する.
主な方法:
- ClinCode Gastro Corpusの800万のノルウェーの臨床ノートで NorDeClin-BERT (ベースと大きい) の2つのバージョンを訓練しました.
- ICD-10の診断コードの予測モデルを 微調整した.
- 精度,精度,リコール,F1スコアを使用して,SweDeClin-BERT,ScandiBERT,NorBERT3-base,NorBERT3-largeと比較したNorDeClin-BERT.
主要な成果:
- NorDeClin-BERTの2つのバージョンは,ICD-10コードの分類において,ノルウェーの一般的なBERTモデルとスウェーデンの臨床BERTモデルを上回った.
- NorDeClin-BERT-largeは,すべての評価指標で最高のパフォーマンスを達成し,分野特有の予備訓練とモデルの能力の利点を示しました.
- スウェーデンの臨床モデルは,ノルウェーに特有の臨床的予備訓練の必要性を強調し,制限された移転性を示しました.
結論:
- NorDeClin-BERTは,ノルウェーの胃腸内科におけるICD-10コードの分類を改善し,ドキュメントの簡素化と行政の負担を軽減する大きな可能性を示しています.
- この研究は,ノルウェーの医療NLPとICD-10のコーディングのための最先端のモデルとして,新しい研究ベースラインを確立しています.
- 将来の研究は,より広範な臨床応用のための高度な領域適応,外部知識統合,および病院間の一般化性を調査する必要があります.
関連する概念動画
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


