Smoothed Nested Testing on Directed Acyclic Graphs
J H Loper1, L Lei2, W Fithian3
1Department of Neuroscience, Columbia University, 716 Jerome L. Greene Building, New York, New York 10025, U.S.A.
Biometrika
|May 2, 2024
Summary
This study introduces a novel smoothing method for multiple hypothesis testing with nested structures. This approach enhances statistical power while controlling error rates, offering significant advantages in complex data analysis.
Area of Science:
- Statistics
- Bioinformatics
- Genomics
Background:
- Multiple hypothesis testing is crucial in analyzing complex datasets.
- Logical nested structures in hypotheses present unique challenges for traditional methods.
- Existing methods may lack power when dealing with hierarchical data relationships.
Purpose of the Study:
- To develop a general framework for hypothesis testing with logical nested structures.
- To propose and evaluate a smoothing procedure to increase statistical power.
- To ensure control of key error rates under various dependency conditions.
Main Methods:
- Modeling hypothesis structures as directed acyclic graphs.
- Adjusting node-level test statistics based on logical constraints.
- Implementing a smoothing procedure combining nodes with descendants.
- Proving error rate control for independent and dependent test statistics.
Main Results:
- A broad class of smoothing strategies effectively controls familywise error rate, false discovery exceedance rate, and false discovery rate.
- Arithmetic averaging demonstrates error rate control even with positively-correlated normal observations.
- Simulations and a biological dataset application show substantial power gains through smoothing.
Conclusions:
- The proposed smoothing framework offers a powerful approach for nested multiple hypothesis testing.
- The method provides robust error rate control across different statistical assumptions.
- This technique has practical implications for biological data analysis and other fields with hierarchical hypotheses.
Related Concept Videos
Wald-Wolfowitz Runs Test I
642
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
642
Wald-Wolfowitz Runs Test II
232
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
232
Types of Hypothesis Testing
26.3K
There are three types of hypothesis tests: right-tailed, left-tailed, and two-tailed.
When the null and alternative hypotheses are stated, it is observed that the null hypothesis is a neutral statement against which the alternative hypothesis is tested. The alternative hypothesis is a claim that instead has a certain direction. If the null hypothesis claims that p = 0.5, the alternative hypothesis would be an opposing statement to this and can be put either p > 0.5, p < 0.5, or p...
When the null and alternative hypotheses are stated, it is observed that the null hypothesis is a neutral statement against which the alternative hypothesis is tested. The alternative hypothesis is a claim that instead has a certain direction. If the null hypothesis claims that p = 0.5, the alternative hypothesis would be an opposing statement to this and can be put either p > 0.5, p < 0.5, or p...
26.3K
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Statistical Hypothesis Testing
1.9K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.9K


