Related Experiment Video
Updated: Jul 22, 2026

14:01
Making Record-efficiency SnS Solar Cells by Thermal Evaporation and Atomic Layer Deposition
Published on: May 22, 2015
[The uncertainties of statistical "significance"]
1Instituto de Ciencias Biomédicas, Facultad de Medicina, Universidad de Chile, Santiago, Chile.
Summary
Statistical significance testing, using P values, is unreliable for scientific discovery. Researchers should report exact P values with effect sizes and confidence intervals instead of declaring "significant" findings.
Area of Science:
- Statistics
- Scientific Methodology
Background:
- Traditional statistical inference, developed by Fisher and Neyman-Pearson, uses P values to determine if observed differences are random or real.
- The common practice involves testing the null hypothesis: if P < 0.05, the difference is deemed "significant"; otherwise, the null hypothesis is accepted.
Purpose of the Study:
- To review the deficiencies of P values in statistical inference.
- To highlight the unreliability of P values for discarding the null hypothesis and for experimental reproducibility.
Main Methods:
- Review of statistical inference principles and common practices.
- Analysis of the concept of statistical significance and its limitations.
- Discussion of the American Statistical Association's recent statement on P values.
Main Results:
- P values are an unreliable measure, with a high probability of incorrectly accepting the null hypothesis.
- Replication of experiments yielding the same P value is improbable.
- The term "significant" is misleading and should be avoided.
Conclusions:
- P values do not adequately measure evidence for a hypothesis.
- Researchers should report exact P values, effect sizes, and confidence intervals.
- Alternative methods to classical significance testing are under investigation.
Related Concept Videos
Statistical Significance
Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
Interpretation of Confidence Intervals
A confidence interval is a better estimate of the population than a point estimate, as it uses a range of values from a sample instead of a single value.
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Errors In Hypothesis Tests
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
Uncertainty: Overview
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
Significance Testing: Overview
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
Accuracy and Errors in Hypothesis Testing
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...

