Related Experiment Video
Updated: Oct 4, 2025

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
Published on: September 20, 2022
Estimation of the proportion of true null hypotheses under sparse dependence: Adaptive FDR controlling in microarray
Aniket Biswas1, Subrata Chakraborty1, Vishwa Jyoti Baruah2
1Department of Statistics, 28675Dibrugarh University, Dibrugarh, Assam, India.
This study introduces a novel clustering method to improve estimates of non-differentially expressed genes in microarray data. The approach enhances adaptive testing procedures, leading to more powerful analysis and identification of significant genes.
Area of Science:
- Bioinformatics
- Statistical Genetics
- Computational Biology
Background:
- Accurate estimation of non-differentially expressed genes is crucial for multiple testing procedures in microarray analysis.
- Existing methods often assume gene expression independence, which is unrealistic given sparse dependence structures in biological data.
- Developing methods to account for sparse dependence is essential for robust gene expression analysis.
Purpose of the Study:
- To propose a novel clustering-based method to estimate the proportion of true null hypotheses in gene expression data.
- To integrate sparse dependence structures into existing estimators for improved accuracy.
- To enhance the power of adaptive multiple testing procedures in microarray data analysis.
Main Methods:
- A clustering approach was developed using a high-dimensional correlation structure under a sparse assumption as a dissimilarity matrix.
- This method was applied to three existing estimators for the proportion of true null hypotheses.
- The proposed method was evaluated through extensive simulations and applied to a colorectal cancer gene expression dataset.
Main Results:
- The proposed clustering method significantly improves existing estimators, making them less conservative.
- The adaptive Benjamini-Hochberg algorithm demonstrated increased power when using the proposed method.
- Application to a colorectal cancer dataset revealed a greater number of identified differentially expressed genes.
Conclusions:
- The novel clustering method effectively accommodates sparse dependence in gene expression data.
- This approach enhances the performance of statistical inference in microarray studies.
- The method offers a valuable tool for identifying differentially expressed genes in complex biological datasets, such as cancer research.
More Related Videos
Related Concept Videos
Identifying Statistically Significant Differences: The F-Test
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Expected Frequencies in Goodness-of-Fit Tests
Statistical Hypothesis Testing
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Chi-square Analysis
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
Fisher's Exact Test

