Related Experiment Video
Updated: Jul 15, 2025

12:18
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.6K
Practical advice on variable selection and reporting using Akaike information criterion.
Chris Sutherland1, Darragh Hare2,3, Paul J Johnson2
1Centre for Research into Ecological and Environmental Modelling, University of St Andrews, St Andrews, UK.
Proceedings. Biological Sciences
|September 27, 2023
Summary
This study clarifies common misunderstandings regarding the Akaike information criterion (AIC) in ecological modeling. It uses simulations to explain
Area of Science:
- Ecology
- Statistics
- Ecological Modeling
Background:
- Model selection is crucial in ecological research, with the Akaike information criterion (AIC) being a dominant tool.
- Common misunderstandings persist among users regarding AIC application, interpretation, and reporting.
- Specific confusion exists around 'pretending' variables and the role of p-values in AIC-based model selection.
Purpose of the Study:
- To address prevalent user misconceptions surrounding the Akaike information criterion (AIC).
- To provide intuitive understanding of AIC application and interpretation through simulation.
- To promote improved statistical practices in ecological model selection and reporting.
Main Methods:
- The study complements existing technical literature on AIC.
- Simulation methods are employed to develop intuition around AIC concepts.
- Focus is placed on interpreting AIC model tables and the relationship between p-values and AIC.
Main Results:
- Simulations provide practical insights into the application of AIC.
- Clarification is offered on the concept of 'pretending' variables in model selection.
- Guidance is provided on the interpretation of statistical support when using AIC.
Conclusions:
- Enhanced understanding of AIC can lead to more robust ecological modeling.
- Simulation-based intuition aids in overcoming common statistical pitfalls.
- The study advocates for better practices in using, interpreting, and reporting AIC-selected models.
Related Concept Videos
Regression Analysis
5.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.8K
Survival Tree
105
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
105
Friedman Two-way Analysis of Variance by Ranks
226
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
226
Variability: Analysis
156
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
156
Factorial Design
13.1K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.1K
Goodness-of-Fit Test
3.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.4K

