電子カルテベースの機械学習予測モデルにおける共変量誤分類の調整
Shuang Yang1, Yonghui Wu1, Mei Liu1
1Department of Health Outcomes and Biomedical Informatics, University of Florida, Gainesville, Florida, USA.
本研究では、電子カルテデータの誤りを修正する手法を導入し、肺がんスクリーニングの予測モデルを改善する。調整されたモデルは、未調整のモデルよりも高い精度を示し、手動レビューなしでバイアスを軽減した。
科学分野:
- バイオメディカルインフォマティクス
- ヘルスサービスリサーチ
- ヘルスケアにおける機械学習
背景:
- 電子カルテ(EHR)には貴重なデータが含まれているが、誤分類エラーが発生しやすい。
- これらのエラーは、予測モデルにバイアスを導入し、臨床的意思決定に影響を与える可能性がある。
- 肺がんスクリーニングへのアドヒアランスのような患者のアウトカムの正確な予測は、非常に重要である。
研究 の 目的:
- EHR由来の共変量の誤分類エラーを調整するための方法論を開発および評価する。
- グループごとおよび個別の重みを使用して、予測モデリングにおけるバイアスを軽減する。
- 肺がんスクリーニングアドヒアランスの予測モデルの精度を向上させる。
主な方法:
- 感度と特異度に基づいて、グループごとおよび個別の重みを使用してEHR共変量を調整する方法論を開発した。
- 肺がんスクリーニングアドヒアランスを予測するために、ロジスティック回帰、XGBoost、ニューラルネットワークを適用した。
- Lung-RADSカテゴリ抽出に自然言語処理(NLP)を利用し、カーネルおよび多項回帰重みを使用して調整した。
- 調整済みモデルと未調整(ナイーブ)モデルおよび真値(オラクル)モデルを、受信者操作特性曲線下面積(AUROC)を使用して比較した。
主要な成果:
- 調整済みモデルは、様々な検証セットサイズ(10%、20%、30%)で、ナイーブモデルを大幅に上回った。
- AUROCの改善は、ナイーブモデルと比較して0.3%から10.4%の範囲であった。
- 調整済みモデルは、オラクルモデルとのパフォーマンスギャップを2.0%-7.5%に縮小した。
- 個別の重みは、グループごとの重みよりも正確なエラー補正を示した。
結論:
- 開発されたフレームワークは、EHR由来の共変量における誤分類バイアスを効果的に軽減する。
- この方法論は、広範な手動データレビューを必要とせずに、肺がんスクリーニングアドヒアランスの予測精度を向上させる。
- このアプローチは、実世界の臨床データを使用した予測モデルの信頼性を向上させるためのスケーラブルなソリューションを提供する。
さらに関連する動画
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
関連する概念動画
Confounding in Epidemiological Studies
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Bias in Epidemiological Studies
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data
