样本分割和评估时间序列的适用性
Richard A Davis1, Leon Fernandes1
1Department of Statistics, Columbia University, 1255 Amsterdam Avenue, New York, New York 10027, USA.
Biometrika
|July 1, 2025
概括
本研究引入了一种用于测试剩余序列独立性的时间序列模型的新型样本分割方法. 这种方法调整了固有的依赖性,改善了适合性评估.
科学领域:
- 统计 统计 统计 统计
- 时间序列分析时间序列分析
- 计量经济学 计量经济学 计量经济学
背景情况:
- 评估时间序列模型合适性通常涉及测试序列独立性的残留物.
- 由于共享的参数估计,拟合余值本质上是依赖的,这复杂化了标准的独立性测试.
- 现有的方法,如样本分割 (Pfister等人,2018) 对于依赖时间序列数据是不够的.
研究的目的:
- 调整样本分割程序,以测试时间序列模型残余中的序列依赖性.
- 为计算剩余依赖的时间序列模型开发调整后的适合性测试.
- 提供一种方法,避免在残余分析中对置信限进行复杂的调整.
主要方法:
- 在时间序列背景下利用样本分割技术.
- 使用数据的初始部分估计模型参数 ([公式:见文本]).
- 使用估计的参数计算后续数据段的剩余值 ([公式:见文本]).
- 将自相关函数 (ACF) 和自距离相关函数 (ADCF) 测试应用于这些计算的余数.
主要成果:
- 拟议的样本分割方法允许 ACF 和 ADCF 测试的序列独立性,以实现相同的极限分布,如果残留物是真正独立的和相同的分布,只要重叠是无意义的.
- 使用数据的前半部分用于参数估计和整个数据集用于剩余计算,可以获得这种理想的属性.
- 这种程序简化了ACF和ADCF在合适性测试中的信心边界的构建.
结论:
- 适应的样本分割技术有效地解决了时间序列模型评估中依赖余量的挑战.
- 这种方法为时间序列模型的合适性测试提供了更可靠的方法.
- 它通过减轻对标准统计测试进行复杂调整的需求来简化剩余分析.
相关概念视频
Goodness-of-Fit Test
4.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.1K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
1.8K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.8K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Survival Tree
166
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
166
Test for Homogeneity
2.1K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.1K
Residuals and Least-Squares Property
7.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.9K


