在非高斯的条件下,因果学习中的共变量选择.
Bixi Zhang1, Wolfgang Wiedermann2
1Department of Educational Psychology, CUNY Graduate Center, New York, NY, USA. bzhang2@gc.cuny.edu.
Behavior research methods
|September 13, 2023
概括
本研究介绍了一种非高斯前选择 (nGFS) 方法,用于在观察性研究中选择控制变量. nGFS算法有效地识别了关键的共变量,改善了因果效应估计,特别是在大样本大小和非高斯数据的情况下.
科学领域:
- 行为科学是一种行为科学.
- 发展科学是一种发展科学.
- 社会科学 社会科学 社会科学
背景情况:
- 了解因果机制是社会和行为科学中的关键.
- 同变量调整对于从观测数据中估计因果关系至关重要.
- 选择合适的共变量是一个挑战,因为单纯的可用性是不够的.
研究的目的:
- 介绍一种用于共变量选择的新型非高斯方法.
- 为线性模型开发一个前向选择算法 (nGFS).
- 通过避免偏见和不一致,提高因果效应估计的准确性.
主要方法:
- 提出了一个非高斯向前选择 (nGFS) 算法.
- 将算法应用于线性模型,用于共变量选择.
- 利用蒙特卡洛模拟研究来评估性能.
主要成果:
- nGFS算法表现良好,特别是在大样本大小 (n ≥ 250) 的情况下.
- 当数据显著偏离高斯度 (倾斜度>1.5) 时,性能会得到提高.
- 该算法与因果模型规范的依赖方向原则保持一致.
结论:
- 在观察性研究中,GGFS方法为共变量选择提供了可靠的方法.
- 这种方法提高了因果效应估计的可靠性.
- 它对于非高斯数据和较大的数据集特别有效.
相关概念视频
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
154
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
154
Causality in Epidemiology
463
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
463
Assumptions of Survival Analysis
153
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
153
Study Design in Statistics
8.2K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
8.2K
Introduction to Nonparametric Statistics
749
Nonparametric statistics offer a powerful alternative to traditional parametric methods, useful when assumptions about the population distribution cannot be made. Unlike parametric tests, which require data to follow a specific distribution with well-defined parameters (such as the mean and standard deviation), nonparametric tests do not require such constraints. This makes them particularly valuable when dealing with small sample sizes, skewed data, or ordinal and categorical variables.
One of...
One of...
749
Randomized Experiments
7.0K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
7.0K


