Related Experiment Video
Updated: Jul 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
An improved approximation to the estimation of the critical F values in best subset regression
David W Salt1, Subhash Ajmani, Ray Crichton
1Department of Mathematics, Buckingham Building, Lion Terrace, University of Portsmouth, Portsmouth, UK.
Abstract:
Variable selection methods are routinely applied in regression modeling to identify a small number of descriptors which "best" explain the variation in the response variable. Most statistical packages that perform regression have some form of stepping algorithm that can be used in this identification process. Unfortunately, when a subset of p variables measured on a sample of n objects are selected from a set of k (>p) to maximize the squared sample multiple regression coefficient, the significance of the resulting regression is upwardly biased. The extent of this bias is investigated by using Monte Carlo simulation and is presented as an inflation factor which when multiplied by the usual tabulated F ratio gives an estimate of the true 5% critical value. The results show that selection bias can be very high even for moderate-size data sets. Selecting three variables from 50 generated at random with 20 observations will almost certainly provide a significant result if the usual tabulated F values are used. An interpolation formula is provided for the calculation of the inflation factor for different combinations of (n, p, k). Four real data sets are examined to illustrate the effect of correlated descriptor variables on the degree of inflation.
Related Concept Videos
Identifying Statistically Significant Differences: The F-Test
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Critical Region, Critical Values and Significance Level
In hypothesis testing, a sample statistic is converted to a test statistic using z, t, or chi-square distribution. A critical region is an area under the curve in probability distributions demarcated by the critical value. When the test statistic falls in this region, it suggests that the null hypothesis must be rejected. As this region contains all those values of the test...
Friedman Two-way Analysis of Variance by Ranks
F Distribution
Expected Frequencies in Goodness-of-Fit Tests
