Related Experiment Video
Updated: Jun 11, 2026

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
Published on: September 20, 2022
Finding distributions that differ, with false discovery rate control
Yonghoon Lee1, Edgar Dobriban1, Eric J Tchetgen Tchetgen1
1Department of Statistics and Data Science, The Wharton School, University of Pennsylvania, 265 South 37th Street, Philadelphia, Pennsylvania 19104, U.S.A.
Abstract:
We consider the problem of comparing a reference distribution with several other distributions. Given a sample from both the reference and the comparison groups, we aim to identify the comparison groups whose distributions differ from that of the reference group. Viewing this as a multiple-testing problem, we introduce a methodology that provides exact, distribution-free control of the false discovery rate. To do so, we introduce the concept of batch conformal p -values and demonstrate that they satisfy positive regression dependence across the groups Benjamini & Yekutieli (2001), thereby enabling control of the false discovery rate through the Benjamini-Hochberg procedure. The proof of positive regression dependence introduces a novel technique for the inductive construction of rank vectors with almost-sure dominance under exchangeability. We evaluate the performance of the proposed procedure through simulations. Despite being distribution-free, in some cases it shows performance comparable to methods with knowledge of the data-generating normal distribution, and it further has more power than direct approaches based on conformal out-of-distribution detection. Furthermore, we illustrate our methods on a hepatitis C treatment dataset, where they identify patient groups with large treatment effects, and on the Current Population Survey dataset, where they identify subpopulations with long working hours.
Related Concept Videos
F Distribution
Identifying Statistically Significant Differences: The F-Test
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used; instead...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...

