Related Experiment Video
Updated: Jul 14, 2025

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
Published on: September 20, 2022
Knockoff procedure for false discovery rate control in high-dimensional data streams
Ka Wai Tsang1, Fugee Tsung2, Zhihao Xu3
1School of Data Science, The Chinese University of Hong Kong, Shenzhen Guangdong 518172, People's Republic of China.
Abstract:
Motivated by applications to root-cause identification of faults in high-dimensional data streams that may have very limited samples after faults are detected, we consider multiple testing in models for multivariate statistical process control (SPC). With quick fault detection, only small portion of data streams being out-of-control (OC) can be assumed. It is a long standing problem to identify those OC data streams while controlling the number of false discoveries. It is challenging due to the limited number of OC samples after the termination of the process when faults are detected. Although several false discovery rate (FDR) controlling methods have been proposed, people may prefer other methods for quick detection. With a recently developed method called Knockoff filtering, we propose a knockoff procedure that can combine with other fault detection methods in the sense that the knockoff procedure does not change the stopping time, but may identify another set of faults to control FDR. A theorem for the FDR control of the proposed procedure is provided. Simulation studies show that the proposed procedure can control FDR while maintaining high power. We also illustrate the performance in an application to semiconductor manufacturing processes that motivated this development.
Related Concept Videos
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Detection of Gross Error: The Q Test
Testing a Claim about Mean: Unknown Population SD
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...
Quantifying and Rejecting Outliers: The Grubbs Test
Wald-Wolfowitz Runs Test II
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
Aliasing
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...

