根据标签缺陷改进基于电子健康记录的临床预测模型:基于网络的生成对抗性半监督方法
Runze Li1, Yu Tian1, Zhuyi Shen1
1College of Biomedical Engineering and Instrument Science, Zhejiang University, Hangzhou, China.
JMIR medical informatics
|June 13, 2023
概括
一种基于网络的创新生成对抗性半监督方法有效地使用有限的电子健康记录 (EHR) 数据训练临床预测模型. 这种方法实现了与监督方法相似的性能,解决了精准医学中数据标签无法访问的问题.
科学领域:
- 生物医学信息学 生物医学信息学
- 医疗保健中的机器学习
- 精准医学是一门精准的医学.
背景情况:
- 电子健康记录 (EHR) 对精准医学至关重要,但数据标签不可访问性阻碍了临床预测.
- 现有的方法,如合成和半监督学习,在利用电子健康记录数据方面存在局限性.
- 需要揭示EHR的底层图形结构,以改善模型培训.
研究的目的:
- 为训练临床预测模型提出基于网络的生成对抗性半监督方法.
- 为了应对标签缺陷的电子健康记录的挑战,并实现与监督方法相比的性能.
- 在不同的基准数据集上评估方法的有效性.
主要方法:
- 使用基于网络的生成对抗性半监督方法.
- 在有5%至25%标记数据的数据集上训练模型.
- 使用分类指标对比传统的半监督和监督方法来评估性能.
- 评估数据质量,模型安全性和内存可扩展性.
主要成果:
- 拟议的半监督方法显著优于现有的半监督方法,实现了高AUC值 (例如,在一个数据集上为0.945).
- 使用10%标记数据的性能与逻辑回归,SVM和随机森林等监督方法相比较.
- 数据综合和隐私保护措施解决了有关数据安全和二次使用的担忧.
结论:
- 训练临床预测模型的标签缺陷的电子健康记录对于数据驱动的研究至关重要.
- 拟议的方法有效地利用了EHR的内在结构,用于稳健的模型开发.
- 这种方法显示了通过改进EHR利用来推进精准医学的巨大潜力.
相关概念视频
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
99
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
99
Methods of Documentation VII: EMR
869
Electronic Medical Records (EMRs) primarily center around electronically documenting patients' health information within a single healthcare organization or practice. They contain essential clinical data related to a patient's medical history, diagnoses, medications, treatment plans, lab results, and other pertinent information relevant to the specific encounter or episode of care. EMRs are designed to streamline documentation and workflow processes within individual healthcare...
869


