Modifying the false discovery rate procedure based on the information theory under arbitrary correlation structure
Sedighe Rastaghi1, Azadeh Saki2, Hamed Tabesh3
1Department of Epidemiology and Biostatistics, School of Health, Mashhad University of Medical Sciences, Mashhad, Iran.
New modified procedures control the False Discovery Rate (FDR) in multiple comparison procedures (MCPs) by accounting for correlations between test statistics. These methods offer improved flexibility and efficiency, especially in high-dimensional data analysis.
Area of Science:
- Statistics
- Bioinformatics
- Genomics
Background:
- Controlling the False Discovery Rate (FDR) is crucial in multiple comparison procedures (MCPs) across scientific fields.
- Test statistic correlation can increase FDR variance and bias, necessitating robust control methods.
- Information theory offers a novel approach to modify MCPs for correlation structures.
Purpose of the Study:
- To develop and evaluate modified MCPs that account for arbitrary correlation structures using information theory.
- To compare the performance of proposed procedures (M1, M2, M3) against established methods like Benjamini-Hochberg (BH) and Benjamini-Yekutieli (BY).
- To assess the utility of these procedures in analyzing high-dimensional gene expression data.
Main Methods:
- Proposed three modified procedures (M1, M2, M3) based on conditional Fisher Information.
- Employed simulation studies with multivariate Gaussian features at varying correlation levels.
- Applied procedures to real high-dimensional colorectal cancer gene expression data.
- Utilized Efficient Bayesian Logistic Regression (EBLR) for predictive modeling.
Main Results:
- Proposed procedures performed similarly to BH when no correlation existed.
- BY was conservative, BH was liberal in low to medium correlations; proposed methods showed reduced feature screening with increased correlation.
- In highly correlated scenarios, proposed methods approached Bonferroni procedure performance.
- EBLR models using features screened by M1 and M2 exhibited minimum entropy and higher efficiency.
Conclusions:
- Modified procedures offer greater flexibility than BH and BY in handling test statistic correlations.
- These information-theory-based methods effectively reduce screening of non-informative features as correlations increase.
- The proposed procedures enhance the efficiency of predictive modeling in high-dimensional data analysis.
More Related Videos
07:11Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
Published on: November 10, 2023
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
Related Concept Videos
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Correlation of Experimental Data
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...
Genetic Variation
Genes exist in different versions called alleles,...
Expected Frequencies in Goodness-of-Fit Tests
Genetic Drift
Mutation, Gene Flow, and Genetic Drift
