Related Experiment Video
Updated: Aug 5, 2025

08:12
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
2.6K
Far from Asymptopia: Unbiased High-Dimensional Inference Cannot Assume Unlimited Data
Michael C Abbott1, Benjamin B Machta1
1Department of Physics, Yale University, New Haven, CT 06520, USA.
Entropy (Basel, Switzerland)
|March 29, 2023
Summary
Jeffreys prior causes bias in high-dimensional models. A new principled measure, focusing on relevant parameters, yields unbiased posteriors, especially with limited data.
Area of Science:
- Statistical modeling
- Information geometry
- Bayesian inference
Background:
- Inference from limited data necessitates a measure on parameter space, often a prior distribution in Bayesian analysis.
- Jeffreys prior, a common uninformative choice derived from information geometry, is shown to introduce significant bias in high-dimensional models.
Purpose of the Study:
- To address the bias introduced by Jeffreys prior in high-dimensional models.
- To propose a principled measure that yields unbiased posteriors by focusing on relevant parameters.
Main Methods:
- Demonstrating the bias of Jeffreys prior in models with effective dimensionality lower than the number of microscopic parameters.
- Developing and presenting a new optimal prior based on relevant parameters and the quantity of data.
Main Results:
- Jeffreys prior exhibits enormous bias in typical high-dimensional scientific models due to unequal treatment of parameters.
- The proposed optimal prior avoids this bias by focusing on relevant parameters, leading to unbiased posteriors.
- This optimal prior is data-dependent and converges to Jeffreys prior only in the asymptotic limit.
Conclusions:
- Standard uninformative priors like Jeffreys prior are inadequate for high-dimensional models common in science.
- A novel, data-dependent prior focusing on relevant parameters offers a principled solution for unbiased Bayesian inference.
- Achieving the asymptotic limit for Jeffreys prior requires an infeasible amount of data in many practical scenarios.
Related Concept Videos
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
173
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
173
Friedman Two-way Analysis of Variance by Ranks
267
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
267
One-Way ANOVA: Unequal Sample Sizes
5.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.8K
One-Way ANOVA: Equal Sample Sizes
3.4K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.4K
Midrange
3.7K
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
3.7K
Introduction to Nonparametric Statistics
795
Nonparametric statistics offer a powerful alternative to traditional parametric methods, useful when assumptions about the population distribution cannot be made. Unlike parametric tests, which require data to follow a specific distribution with well-defined parameters (such as the mean and standard deviation), nonparametric tests do not require such constraints. This makes them particularly valuable when dealing with small sample sizes, skewed data, or ordinal and categorical variables.
One of...
One of...
795

