时间序列数据的高维模拟推理
Chien-Ming Chi1, Yingying Fan2, Ching-Kang Ing3
1Institute of Statistical Science, Academia Sinica, Taiwan.
Journal of the American Statistical Association
|August 26, 2025
概括
我们介绍了时间序列对应推断 (TSKI),这是一个用于在时间序列数据中进行强有力的特征选择的新方法. TSKI解决了序列依赖和未知的共变量分布,控制了错误发现率 (FDR).
科学领域:
- 统计数据
- 时间序列分析
- 机器学习
背景情况:
- 模型X仿制推断是功能选择的强大工具,但由于串行依赖,它面临时间序列数据的挑战.
- 现有的方法通常需要严格的假设关于共变量分布,这对于现实世界时间序列来说是不切实际的.
- 解决这些局限性对于动态系统中可靠的特征选择至关重要.
研究的目的:
- 建立专门针对时间序列数据的仿制推理的理论和方法基础.
- 开发一种新的方法,即时间序列模仿推断 (TSKI),克服现有方法的局限性.
- 通过在具有挑战性的时间序列条件下控制错误发现率 (FDR) 来确保强大的特征选择.
主要方法:
- 通过整合子样本和e值来管理序列依赖性,提出时间序列对应推断 (TSKI).
- 一般化强大的仿制推断,放松已知的共变量分布的假设,使其适用于时间序列.
- 建立了对非对称错误发现率 (FDR) 的理论条件,并使用拉索进行了功率分析.
主要成果:
- 证明TSKI在足够的条件下有效控制非对称错误发现率 (FDR).
- 通过技术分析量化了序列依赖和未知的共变量分布对FDR控制的影响.
- 通过模拟和经济通胀研究验证了TSKI的有限样本表现.
结论:
- TSKI为时间序列数据中的特征选择提供了强大且理论上可靠的框架.
- 该方法成功地解决了序列依赖和未知的共变量分布的复杂性.
- 根据经验评估,TSKI提供了可靠的时间序列分析推断的实用解决方案.
相关概念视频
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
208
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
208
Quantifying and Rejecting Outliers: The Grubbs Test
2.0K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.0K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
710
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
710
Time-Series Graph
4.5K
A time-series graph is a line graph with repeated measurements taken at successive intervals of time. It is also called a time series chart. To construct a time-series graph, one must look at both pieces of a paired data set. The horizontal axis is used to plot the time increments, and the vertical axis is used to plot the values of the variable that one is measuring. By using the axes in this way, each point on the graph will correspond to time and a measured quantity. The points on the graph...
4.5K
Outliers and Influential Points
4.2K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.2K
Friedman Two-way Analysis of Variance by Ranks
296
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
296


