利用易出错的算法衍生的表型:加强对电子健康记录数据中的风险因素的关联研究
Yiwen Lu1, Jiayi Tong2, Jessica Chubak3
1Center for Health AI and Synthesis of Evidence (CHASE), Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA; The Graduate Group in Applied Mathematics and Computational Science, School of Arts and Sciences, University of Pennsylvania, Philadelphia, PA, USA.
这项研究引入了一种新方法,将多个电子健康记录 (EHR) 现型结合起来,减少偏见并提高关联研究的效率. 该方法提高了统计准确性,并为表型/暴露分析提供了强大的替代方案.
科学领域:
- 生物医学信息学 生物医学信息学
- 健康 数据科学 数据科学
- 临床流行病学临床流行病学
背景情况:
- 多个可计算的表型越来越多地来自电子健康记录 (EHR).
- 基于EHR的关联研究通常使用单个表型,可能引入偏见.
- 表型错误可能会影响基于EHR的研究的准确性和效率.
研究的目的:
- 开发一种用于同时利用多个EHR衍生的表型的新方法.
- 为了减少EHR数据中的表型错误引起的偏见.
- 提高表型/暴露关联研究的效率和准确性.
主要方法:
- 开发了一种将多个算法衍生的表型与验证的结果结合在一起的方法.
- 该方法采用了统计学上高效的看似无关的回归框架.
- 通过模拟研究和现实世界EHR数据分析 (结肠癌复发) 来评估性能.
主要成果:
- 与单个表型方法相比,实现了实质性的偏差减少,特别是当没有单个表型均优越时.
- 与仅使用一个算法衍生的表型相比,估计效率提高了多达30%.
- 该方法在整合多种表型以提高统计准确度方面表现出有效性.
结论:
- 拟议的方法有效地整合了来自EHR数据的多种表型.
- 它提供了一个强大的替代单替代偏差校正方法.
- 该方法提高了基于EHR的关联研究中的偏差减少,统计准确性和效率.
更多相关视频
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
05:53Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
相关概念视频
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Methods of Documentation VII: EMR
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Confounding in Epidemiological Studies
Epistasis Analysis
