使用仅预测缺失的辅助归算变量进行多次归算,可能会增加因数据缺失而导致的偏差,而不是随机的数据缺失
Elinor Curnow1,2, Rosie P Cornish3,4, Jon E Heron3,4
1Department of Population Health Sciences, Bristol Medical School, University of Bristol, Bristol, UK. elinor.curnow@bristol.ac.uk.
BMC medical research methodology
|October 7, 2024
概括
在多重归算 (MI) 中包括不相关的辅助变量可以在数据缺失时加剧偏差,而不是随机 (MNAR). 研究人员应该根据他们对缺失数据的预测能力,仔细选择辅助变量,而不仅仅包括所有可用的数据.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
背景情况:
- 流行病学和临床研究经常遇到缺失的数据,通常使用多重归咎 (MI) 来处理.
- 如果数据不随机 (MNAR) 缺失,多重归算估计可能会有偏见.
- 辅助变量可以减轻MNAR数据的MI偏差,但选择策略需要仔细考虑.
研究的目的:
- 探索包括辅助变量预测缺失的影响,但与MI模型中部分观察到的变量无关.
- 量化线性或逻辑回归模型中由这些辅助变量引入的额外偏差.
- 评估对暴露系数的影响,当部分观察到的变量是结果或暴露时.
主要方法:
- 使用代数量化和模拟研究来评估偏差.
- 这项研究比较了含有和不含特定类型辅助变量的MI模型.
- 分析涉及连续和二进制部分观察变量,应用于结果和暴露.
- 从出生队列研究中重新分析数据说明了这些发现.
主要成果:
- 包括一个预测缺失的辅助变量,但与部分观察到的变量无关,可以引入实质性的额外偏差.
- 这种额外的偏差特别明显,当结果被部分观察时,缺失是由结果本身驱动的.
- 当结果和暴露都对失踪机制有所贡献时,偏见甚至更大.
结论:
- 在数据可能是MNAR时,应避免在MI中包含所有可用的辅助变量的常见做法.
- 辅助变量应根据它们与部分观察到的变量有很强的预测关联来选择.
- 确定适当的辅助变量需要仔细考虑因果图,缺失机制和数据探索,考虑潜在的选择偏差.
相关概念视频
Bias
3.8K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.8K
Biostatistics: Overview
227
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
227
Bias in Epidemiological Studies
167
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
167
Assumptions of Survival Analysis
101
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
101
Random Error
843
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
843
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
117
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
117


