Related Experiment Video
Updated: Jun 11, 2026

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
Published on: September 20, 2022
Finding distributions that differ, with false discovery rate control
Yonghoon Lee1, Edgar Dobriban1, Eric J Tchetgen Tchetgen1
1Department of Statistics and Data Science, The Wharton School, University of Pennsylvania, 265 South 37th Street, Philadelphia, Pennsylvania 19104, U.S.A.
This study introduces batch conformal p-values for comparing distributions, offering exact, distribution-free false discovery rate control. This method identifies differing groups and shows strong performance in simulations and real-world datasets.
Area of Science:
- Statistics
- Machine Learning
Background:
- Comparing multiple distributions to a reference is crucial in data analysis.
- Existing methods may lack robustness or require distributional assumptions.
Purpose of the Study:
- To develop a novel methodology for identifying comparison groups with distributions differing from a reference group.
- To provide exact, distribution-free control of the false discovery rate (FDR) in multiple-testing scenarios.
Main Methods:
- Introduction of 'batch conformal p-values'.
- Demonstration of positive regression dependence across groups, enabling FDR control via the Benjamini-Hochberg procedure.
- Novel proof technique for rank vector construction under exchangeability.
Main Results:
- The proposed method achieves exact, distribution-free FDR control.
- Simulations show performance comparable to distribution-specific methods and superior power over direct conformal detection.
- Application to hepatitis C and Current Population Survey datasets identified significant patient and subpopulation groups.
Conclusions:
- Batch conformal p-values offer a powerful, distribution-free approach for multiple distribution comparison.
- The methodology is effective in identifying meaningful differences in real-world data.
- This work advances FDR control techniques in statistical inference.
Related Concept Videos
F Distribution
Identifying Statistically Significant Differences: The F-Test
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used; instead...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...

