Related Experiment Video
Updated: May 5, 2026

10:10
Three Differential Expression Analysis Methods for RNA Sequencing: limma, EdgeR, DESeq2
Published on: September 18, 2021
42.7K
A comparison of two classes of methods for estimating false discovery rates in microarray studies
Emily Hansen1, Kathleen F Kerr
1Cancer Research and Biostatistics, Seattle, WA 98101, USA.
Scientifica
|November 27, 2013
Summary
Estimating the proportion of differentially expressed genes (π 1) is key for microarray analysis. Directly using test statistics is less reliable than using P-values with mixture models for accurate false discovery rate control.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Genetics
Background:
- Microarray studies aim to identify differentially expressed genes between populations.
- Estimating the false discovery rate (FDR) is crucial for interpreting these gene lists.
- Accurate FDR estimation relies on estimating π 1, the proportion of truly differentially expressed genes.
Purpose of the Study:
- To evaluate methods for estimating π 1 directly from test statistics, bypassing P-value computation.
- To compare the performance of these novel methods against established P-value-based approaches.
- To identify the most reliable methods for estimating π 1 in microarray data analysis.
Main Methods:
- Adapted existing methods to estimate π 1 from t- and z-statistics.
- Developed methods for direct π 1 estimation from various test statistics.
- Compared these new methods against two established π 1 estimation techniques using P-values.
Main Results:
- Methods for estimating π 1 varied significantly in bias and variability.
- Estimates derived directly from test statistics did not perform reliably.
- The 'convest' mixture model applied to P-values from a pooled permutation null distribution yielded the least biased and most variable estimates of π 1.
Conclusions:
- Computing P-values, while sometimes viewed as problematic, remains a necessary step for reliable π 1 estimation.
- The 'convest' method with permutation-based P-values offers superior accuracy for controlling FDR in gene expression studies.
- Direct estimation of π 1 from test statistics is not recommended based on current findings.
Related Concept Videos
DNA Microarrays
16.9K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
16.9K
Multiple Comparison Tests
3.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.5K
Comparing Copy Number Variations and SNPs
11.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.6K
Bonferroni Test
2.6K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.6K
Comparing the Survival Analysis of Two or More Groups
726
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
726
Testing a Claim about Population Proportion
2.9K
A complete procedure for testing a claim about a population proportion is provided here.
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
2.9K

