Sample splitting and assessing goodness-of-fit of time series
Richard A Davis1, Leon Fernandes1
1Department of Statistics, Columbia University, 1255 Amsterdam Avenue, New York, New York 10027, USA.
Biometrika
|July 1, 2025
Summary
This study introduces a novel sample-splitting method for time series models to test residual serial independence. This approach adjusts for inherent dependencies, improving goodness-of-fit assessments.
Area of Science:
- Statistics
- Time Series Analysis
- Econometrics
Background:
- Assessing time series model fit often involves testing residuals for serial independence.
- Fitted residuals are intrinsically dependent due to shared parameter estimates, complicating standard independence tests.
- Existing methods like sample splitting (Pfister et al., 2018) are insufficient for dependent time series data.
Purpose of the Study:
- To adapt the sample-splitting procedure for testing serial dependence in time series model residuals.
- To develop adjusted goodness-of-fit tests for time series models that account for residual dependence.
- To provide a method that avoids complex adjustments for confidence bounds in residual analysis.
Main Methods:
- Leveraging a sample-splitting technique in the time series context.
- Estimating model parameters using an initial segment of the data ([Formula: see text]).
- Computing residuals for a subsequent data segment ([Formula: see text]) using the estimated parameters.
- Applying autocorrelation function (ACF) and auto-distance correlation function (ADCF) tests to these computed residuals.
Main Results:
- The proposed sample-splitting method allows ACF and ADCF tests of serial independence to achieve the same limit distributions as if residuals were truly independent and identically distributed, provided overlap is asymptotically negligible.
- Using the first half of the data for parameter estimation and the entire dataset for residual computation yields this desirable property.
- This procedure simplifies the construction of confidence bounds for ACF and ADCF in goodness-of-fit testing.
Conclusions:
- The adapted sample-splitting technique effectively addresses the challenge of dependent residuals in time series model evaluation.
- This method offers a more reliable approach to goodness-of-fit testing for time series models.
- It simplifies residual analysis by mitigating the need for complex adjustments to standard statistical tests.
Keywords:
AutocorrelationAutoregressive moving average modelDistance covarianceGoodness-of-fit testSample splittingTime seriesgarch modelMore Related Videos
Related Concept Videos
Goodness-of-Fit Test
4.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
4.1K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
1.8K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.8K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Survival Tree
166
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
166
Test for Homogeneity
2.1K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.1K
Residuals and Least-Squares Property
7.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.9K


