在缺失混因子的观测数据中评估治疗效果:对实用双强和传统缺失数据方法的比较研究
Brian D Williamson1,2, Chloe Krakauer1, Eric Johnson1
1Biostatistics Division, Kaiser Permanente Washington Health Research Institute, Seattle, Washington, USA.
Statistics in medicine
|February 4, 2026
概括
双强度的方法可以提高在药物流行病学研究中处理缺少数据的效率. 这些未被充分利用的技术,一般化和有针对性的最大概率估计 (TMLE),优于传统方法,如多重归算 (MI) 和反向概率权重 (IPW).
科学领域:
- 药理流行病学 药理流行病学
- 生物统计学 生物统计学
- 健康 数据科学 数据科学
背景情况:
- 药物流行病学依赖于行政和电子健康记录 (EHR) 进行安全性和有效性研究.
- 在这些真实世界数据集中,缺少混数据是常见的,需要强大的分析方法.
- 多重归算 (MI) 和逆概率权重 (IPW) 是标准的,但对于缺少的数据可能不是最佳的.
研究的目的:
- 评估和比较使用不足的双重强大的方法与传统MI和IPW的性能,以处理药物流行病学中缺少的数据.
- 根据各种场景的性能,为选择适当的缺失数据分析方法提供指导.
主要方法:
- 调查了两个双重强大的估计器:通用拉开和反向概率加权的目标最大概率估计 (TMLE).
- 通过综合数据进行了广泛的数值研究,涉及各种失踪情况和数据生成场景,包括罕见的结果和高失踪比例.
- 利用模拟大型EHR队列数据的等离子体模拟研究来评估在罕见的结果设置中与大量缺失的混数据 (> 50%) 的性能.
主要成果:
- 与MI和IPW相比,双倍强大的方法在各种场景中表现出优异的性能,特别是在偏差差异权衡方面,与MI和IPW相比.
- 一般化和TMLE显示出更高的效率和稳定性,特别是在高缺失率和罕见结果的情况下.
- 该研究确定了特定的场景,其中双重强度的方法显著减少偏差,提高效果估计的精度.
结论:
- 双重强大的方法,特别是通用和TMLE,代表了有价值的,但尚未充分利用的,用于在药理流行病学中缺少数据分析的工具.
- 这些方法提供了更好的统计效率和稳定性,从而从现实数据中获得更可靠的安全性和有效性估计.
- 鼓励研究人员采用双重强大的方法来处理缺少的混数据,以提高药物流行病学发现的有效性.
相关概念视频
Data Collection by Observations
15.0K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
15.0K
Assessment of the Gastrointestinal System I: Subjective Data
671
Assessing the gastrointestinal (GI) system is a complex process that begins with collecting subjective data. This data, collected through patient interviews, provides crucial insights into the patient's health history, perception patterns, and lifestyle habits, all contributing significantly to GI health.
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems...
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems...
671
Assessment of the Cardiovascular System I: Subjective Data
855
A thorough health history and physical assessment are essential for identifying cardiovascular disease (CVD) symptoms and distinguishing them from other health issues.
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
855
How Data are Classified: Categorical Data
44.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.8K
Confounding in Epidemiological Studies
842
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
842
How Data are Classified: Numerical Data
38.1K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
38.1K


