Related Experiment Video
Updated: Apr 28, 2026

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
Published on: September 20, 2022
Variable Selection with FDR Control for Noisy Data - An Application to Screening Metabolites that Are Associated with
Runqiu Wang1, Ran Dai1, Ying Huang2
1Department of Biostatistics, University of Nebraska Medical Center, Omaha, Nebraska, U.S.A.
None:
The rapidly expanding field of metabolomics presents an invaluable resource for understanding the associations between metabolites and various diseases. However, the high dimensionality, presence of missing values, and measurement errors associated with metabolomics data can present challenges in developing reliable and reproducible approaches for disease association studies. Therefore, there is a compelling need for robust statistical analyses that can navigate these complexities to achieve reliable and reproducible disease association studies. In this paper, we construct algorithms to perform variable selection for noisy data and control the False Discovery Rate when selecting mutual metabolomic predictors for multiple disease outcomes. We illustrate the versatility and performance of this procedure in a variety of scenarios, dealing with missing data and measurement errors. As a specific application of this novel methodology, we target two of the most prevalent cancers among US women: breast cancer and colorectal cancer. By applying our method to the Women's Health Initiative data, we successfully identify metabolites that are associated with either or both of these cancers, demonstrating the practical utility and potential of our method in identifying consistent risk factors and understanding shared mechanisms between diseases.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Frequency-dependent Selection
Censoring Survival Data
Detection of Gross Error: The Q Test
Randomized Experiments
Simple randomization
Simple...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...

