Related Experiment Video
Updated: Oct 3, 2025

13:55
Combined Immunofluorescence and DNA FISH on 3D-preserved Interphase Nuclei to Study Changes in 3D Nuclear Organization
Published on: February 3, 2013
18.6K
The generalized Fisher's combination and accurate p-value calculation under dependence.
1Biostatistics and Research Decision Sciences, Merck Research Laboratories, Rahway, New Jersey, USA.
Biometrics
|February 18, 2022
Summary
Accurate p-value calculations are crucial for combining dependent significance tests. This study introduces the GFisher statistic and novel methods to improve accuracy and reduce false discoveries in big data analysis.
Area of Science:
- Statistics
- Bioinformatics
- Computational Biology
Background:
- Combining dependent significance tests is vital in big data analysis.
- Existing methods like Fisher's combination test struggle with accurate p-value calculation, especially at low significance levels, leading to inflated Type I error rates and false discoveries.
- This issue is particularly problematic in large-scale datasets where even small error rates can yield significant false positives.
Purpose of the Study:
- To develop a generalized Fisher-type statistic (GFisher) encompassing various combination methods.
- To introduce novel p-value calculation techniques for improved accuracy and robustness.
- To address the challenge of inflated Type I error rates in significance testing for big data.
Main Methods:
- Introduced the generalized Fisher (GFisher) statistic, a flexible framework for combining dependent tests.
- Developed new p-value calculation methods using moment-ratio matching and joint-distribution surrogating.
- Validated methods through systematic simulations under various distributions (multivariate Gaussian, generalized linear model, multivariate t-distribution).
Main Results:
- The GFisher framework integrates multiple classic statistics with adaptable weighting schemes.
- New p-value calculation methods demonstrate superior accuracy and robustness compared to existing approximations.
- Simulations confirm improved performance, particularly under non-ideal distributional assumptions.
Conclusions:
- The GFisher statistic and new p-value calculation methods offer a more reliable approach for combining dependent significance tests.
- These advancements can mitigate false discoveries in big data research, enhancing analytical rigor.
- The GFisher R package is available for practical application, including SNP-set association studies.
Related Concept Videos
Fisher's Exact Test
854
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of...
854
Behrens–Fisher Test
144
The Behrens-Fisher test is a statistical method designed to address the Behrens-Fisher problem, which arises when comparing the means of two normally distributed populations with unequal variances. Unlike the Student's t-test, which assumes equal variances, the Behrens-Fisher test allows for mean comparison without this restrictive assumption. This flexibility makes it particularly valuable in scenarios where two independent samples exhibit normality but lack variance homogeneity.
This test...
This test...
144
Introduction to Test of Independence
2.5K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.5K
Decision Making: P-value Method
5.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.8K
Determination of Expected Frequency
2.3K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.3K
Bonferroni Test
2.9K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.9K

