在高维数据流中控制错误发现率的淘汰程序
Ka Wai Tsang1, Fugee Tsung2, Zhihao Xu3
1School of Data Science, The Chinese University of Hong Kong, Shenzhen Guangdong 518172, People's Republic of China.
Journal of applied statistics
|October 9, 2023
概括
本研究引入了一种新的Knockoff过程序,用于识别统计过程控制 (SPC) 中错误的数据流. 该方法有效控制错误发现,同时保持高功率,即使在有限的失控样本.
科学领域:
- 统计 统计 统计 统计
- 工业工程 工业工程 工业工程
- 数据科学数据科学数据科学
背景情况:
- 在高维数据流中的根源原因识别对于故障检测至关重要.
- 控制之外 (OC) 数据流的有限样本对传统方法构成挑战.
- 控制错误发现对于统计过程控制 (SPC) 中可靠的故障识别至关重要.
研究的目的:
- 在多变量SPC中为多重测试提出一种新的Knockoff程序.
- 开发一种与现有的故障检测技术相结合的方法,而不会改变停止时间.
- 在识别OC数据流时控制错误发现率 (FDR).
主要方法:
- 对于多变量SPC,建议使用Knockoff过程序.
- 该程序旨在与其他故障检测方法集成.
- 通过定理提供FDR控制的理论保证.
主要成果:
- 拟议的Knockoff程序有效控制了错误发现率 (FDR).
- 该方法在识别有缺陷的数据流方面表现出高的统计能力.
- 模拟研究验证了性能和FDR控制.
结论:
- 淘汰程序提供了一个可靠的解决方案,用于在有限的OC样本中识别SPC中的故障.
- 这种方法提高了高维数据流中故障检测的可靠性.
- 该方法适用于现实世界的场景,例如半导体制造.
相关概念视频
Downsampling
169
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
169
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Testing a Claim about Mean: Unknown Population SD
3.5K
A complete procedure of testing a hypothesis about a population mean when the population standard deviation is unknown is explained here.
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used;...
3.5K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Wald-Wolfowitz Runs Test II
255
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
255
Aliasing
146
Accurate signal sampling and reconstruction are crucial in various signal-processing applications. A time-domain signal's spectrum can be revealed using its Fourier transform. When this signal is sampled at a specific frequency, it results in multiple scaled replicas of the original spectrum in the frequency domain. The spacing of these replicas is determined by the sampling frequency.
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
146


