Related Experiment Video
Updated: Jul 18, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
Accurate and fast small p-value estimation for permutation tests in high-throughput genomic data analysis with the
Yang Shi1,2,3, Weiping Shi4, Mengqiao Wang5
1Division of Biostatistics and Data Science, Department of Population Health Sciences and Department of Neuroscience and Regenerative Medicine, Medical College of Georgia, Augusta University, Augusta, GA 30912, USA.
New algorithms efficiently estimate small p-values in permutation tests for genomic data. This significantly reduces computational effort, improving analysis speed for gene expression studies.
Area of Science:
- Genomics
- Computational Biology
- Statistical Genetics
Background:
- Permutation tests are crucial for hypothesis testing when distributions are complex.
- Estimating very small p-values in genomic studies requires extensive computation.
- Existing methods face computational challenges with large genomic datasets.
Purpose of the Study:
- To develop accurate and efficient algorithms for estimating small p-values in permutation tests.
- To address the computational intensity of permutation tests in genomic data analysis.
- To provide improved solutions for hypothesis testing in genomics.
Main Methods:
- Developed novel algorithms for paired and independent two-group genomic data.
- Utilized Bernoulli and conditional Bernoulli distributions to parameterize sample spaces.
- Employed the cross-entropy method for efficient estimation.
- Leveraged novel frameworks for parameterizing permutation sample spaces.
Main Results:
- Achieved orders of magnitude computational efficiency gains in estimating small p-values.
- Demonstrated performance on simulated and real-world gene expression datasets (microarray, RNA-Seq).
- Outperformed existing methods like crude permutations and SAMC in efficiency.
Conclusions:
- Proposed algorithms offer significant computational advantages for genomic permutation tests.
- These methods enhance the efficiency of existing permutation test procedures.
- The framework facilitates the development of new permutation-based genomic analysis tools.
Related Concept Videos
P-value
P-value stands for the probability value. P-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme as the results obtained from the given sample.
A large P-value calculated from the data indicates to not reject the null hypothesis. But a higher P-value does not mean that the null hypothesis is true. The smaller the P-value, the more...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Chi-square Analysis
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
Wald-Wolfowitz Runs Test II
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
Hardy-Weinberg Principle
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...

