从电子健康记录中描述和分析部分观察到的混数据的原则方法
Janick Weberpals1, Sudha R Raman2, Pamela A Shaw3
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Clinical epidemiology
|May 27, 2024
概括
了解电子健康记录 (EHR) 中缺失的数据至关重要. 这项研究开发了诊断来识别失踪机制,并发现多重归算有效地减少了EHR分析中的偏差.
科学领域:
- 生物统计学 生物统计学
- 医疗信息学 医疗信息学
- 数据科学数据科学数据科学
背景情况:
- 在电子健康记录 (EHR) 中部分观察到的混数据存在重大分析挑战.
- 经常缺乏对基本缺失数据机制的系统评估,这阻碍了可靠的统计分析.
研究的目的:
- 开发一种原则性的方法,以实证地描述EHR中缺少的数据流程.
- 在处理缺失的混数据时,调查各种分析方法的性能.
主要方法:
- 模拟缺失数据机制 (MCAR,MAR,MNAR) 对于使用等离子模式框架的糖尿病患者队列中感兴趣的混因子.
- 基于患者特征差异 (ASMD),缺失指标预测和结果关联的评估诊断.
- 进行比较分析方法,包括完整案例分析,反向概率加权和单/多次归算.
主要成果:
- 经验诊断成功地确定了不同失踪机制 (MAR,MNAR) 的独特模式.
- 不随机失踪 (MNAR) 与结果显著相关,与MAR或MCAR不同.
- 使用随机森林算法的多重归算显示出最低的根-平均-平方-误差,表明性能优越.
结论:
- 开发的诊断提供了对EHR中缺失的数据机制的可靠见解.
- 多重归算,特别是在非参数模型中,可以有效地减少假设满足时的偏差.
相关概念视频
Strategies for Assessing and Addressing Confounding
93
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
93
Confounding in Epidemiological Studies
164
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
164
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
60
Noncompartmental analyses offer an alternative method for describing drug pharmacokinetics without relying on a specific compartmental model. In this approach, the drug's pharmacokinetics are assumed to be linear, with the terminal phase log-linear. This assumption allows for simplified analysis and interpretation of the drug's behavior in the body.
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
60
Bias in Epidemiological Studies
246
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
246
Statistical Methods for Analyzing Epidemiological Data
361
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
361
Mechanistic Models: Compartment Models in Individual and Population Analysis
37
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
37


