在仪表变量存在的情况下,在通用线性模型中测试遗漏的随机假设
Rui Duan1, C Jason Liang2, Pamela A Shaw3,4
1Department of Biostatistics, Harvard T. H. Chan School of Public Health, Boston, Massachusetts, USA.
Scandinavian journal of statistics, theory and applications
|February 19, 2024
概括
本研究引入了一种新的假设测试方法,以区分在通用线性模型中随机缺失和非随机缺失的数据机制. 该方法提供了一个客观的,数据驱动的方式,在处理缺失数据时选择适当的统计程序.
科学领域:
- 统计 统计 统计 统计
- 生物统计学 生物统计学
- 计量经济学 计量经济学
背景情况:
- 数据缺失是统计分析中普遍存在的问题,影响研究结果的有效性和效率.
- 了解数据缺失的机制 (随机缺失与非随机缺失) 对于准确的统计推断至关重要.
- 现有的方法往往难以明确区分这些缺失数据机制.
研究的目的:
- 开发一种新的假设测试框架,以区分随机缺失 (MAR) 和不随机缺失 (MNAR) 数据.
- 提供以数据为导向的客观方法,用于在通用线性模型 (GLM) 中选择正确的缺失数据机制.
- 纳入工具变量以帮助识别缺失数据机制.
主要方法:
- 提出了一种新的假设测试方法,该方法基于估计者之间的差异测量.
- 这些估计器表现出不同的特性,特别是当缺失的数据机制不是随机缺失时.
- 该方法应用于带有仪器变量的通用线性模型的背景下.
主要成果:
- 开发的测试方法提供了一个客观的,数据驱动的MAR和MNAR之间的决定.
- 理论分析证实了拟议测试方法的有效性和有效性.
- 模拟研究和真实数据分析证明了该方法的实际可行性.
结论:
- 新的假设测试方法提供了一个强大的解决方案,用于识别GLM中缺失的数据机制.
- 这有助于在处理缺失数据时更适当的统计分析和可靠的结论.
- 该方法通过严格的理论,模拟和经验证据来验证.
相关概念视频
Assumptions of Survival Analysis
127
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
127
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Errors In Hypothesis Tests
4.2K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.2K
Randomized Experiments
6.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.9K
Friedman Two-way Analysis of Variance by Ranks
196
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
196


