Related Experiment Video
Updated: Feb 5, 2026

13:19
Microarray-based Identification of Individual HERV Loci Expression: Application to Biomarker Discovery in Prostate Cancer
Published on: November 2, 2013
17.2K
Using controls to limit false discovery in the era of big data
Matthew M Parks1, Benjamin J Raphael2, Charles E Lawrence3,4
1Department of Physiology and Biophysics, Weill Cornell Medicine, 1300 York Ave, New York, NY, 10065, USA.
BMC Bioinformatics
|September 16, 2018
Summary
This study introduces a novel method for controlling the false discovery rate (FDR) in big data by using control data to empirically build distributions, avoiding problematic p-value calculations and extrapolation.
Area of Science:
- Statistics
- Bioinformatics
- Genomics
Background:
- Traditional false discovery rate (FDR) control methods in high-dimensional statistics rely on accurate p-values and extrapolation of null distributions.
- These intermediate steps can be challenging and compromise the reliability of statistical results.
- The multiple comparisons problem is a significant concern in big data analysis.
Purpose of the Study:
- To develop a general FDR controlling method suitable for big data settings.
- To overcome limitations of current FDR procedures, specifically the reliance on p-values and distribution extrapolation.
- To leverage abundant control data in big data studies for more robust inference.
Main Methods:
- A novel FDR control method is presented that uses control data to empirically construct the null distribution.
- The empirical null distribution is directly compared to the empirical test data distribution.
- The method avoids the calculation of p-values and extrapolation into unknown distribution tails.
Main Results:
- The proposed control data-based empirical FDR procedure provides a more direct comparison between test and control data.
- This approach aligns with fundamental scientific principles of inference.
- The method is successfully demonstrated in a structural genomics application.
Conclusions:
- A general statistical framework for FDR control tailored to big data is established.
- The method bypasses potentially problematic modeling and extrapolation steps by using empirical distributions and control data.
- The procedure is broadly applicable where controlled experiments or internal negative controls are available, common in big data research.
Related Concept Videos
False Memories
481
False memories represent a cognitive distortion in which individuals recall events that did not happen, or remember them in an altered form. This phenomenon highlights the brain's constructive nature in processing and recalling memories, emphasizing that memory is not a perfect representation of past events but rather a dynamic reconstruction influenced by various factors.
One primary source of false memories is misattribution, where individuals incorrectly associate external information...
One primary source of false memories is misattribution, where individuals incorrectly associate external information...
481
Limiting Reactant
70.1K
The relative amounts of reactants and products represented in a balanced chemical equation are often referred to as stoichiometric amounts. However, in reality, the reactants are not always present in the stoichiometric amounts indicated by the balanced equation.
70.1K
The Number e as a Limit
91
The number e is a fundamental constant in calculus, playing a central role in describing continuous change, particularly exponential growth. It is most naturally defined through its relationship with the natural logarithm, which is the inverse of the exponential function with base e. This relationship allows e to be characterized using basic principles of differentiation rather than as an arbitrary numerical constant.A key property of the natural logarithm function, ln x, is that its derivative...
91
Drug Discovery: Overview
11.6K
Drug discovery is a multifaceted process involving extensive screening, testing, and optimization of lead compounds to identify potential new drugs for therapeutic use. It combines several approaches, including screening large numbers of natural products, chemical modification of known active molecules, identification of new drug targets, and rational design based on biological mechanisms and drug-receptor structure. These approaches are carried out in both academic research laboratories and...
11.6K
Types of Limits I
190
Limits are a key mathematical concept for understanding how functions behave as their input approaches specific values, particularly when the function is undefined. They help reveal trends and discontinuities by examining the values a function approaches rather than its actual value.One-sided limits focus on the direction from which a value is approached. When a function behaves differently depending on whether the input approaches from the left or the right, the two one-sided limits may not...
190
Limit Laws I
228
Limit laws provide essential tools for analyzing how functions behave as their input approaches a specific value. These laws are particularly useful when dealing with combinations of functions, provided the individual limits exist. The Sum and Difference Laws state that the limit of the sum or difference of two functions equals the sum or difference of their respective limits:The Product Law asserts that the limit of the product of two functions equals the product of their individual limits:A...
228

