Related Experiment Video
Updated: May 9, 2025

14:34
How to Create and Use Binocular Rivalry
Published on: November 10, 2010
75.1K
Evaluating the median p-value method for assessing the statistical significance of tests when using multiple
Peter C Austin1,2,3, Iris Eekhout4, Stef van Buuren4,5
1ICES, Toronto, Canada.
Journal of Applied Statistics
|April 30, 2025
Summary
The median p-value method inflates statistical significance when pooling test statistics across imputed datasets. This method should not be used due to unreliable results, especially with increased missing data.
Area of Science:
- Statistics
- Biostatistics
- Data Science
Background:
- Multiple imputation is a standard technique for handling missing data in statistical analyses.
- Rubin's Rules are widely used for pooling results from multiply imputed datasets but are not suitable for test statistics.
- Existing methods for pooling test statistics are complex and not widely implemented in software.
Purpose of the Study:
- To evaluate the performance of the median p-value method for pooling test statistics across imputed samples.
- To determine the reliability of the median p-value method for assessing statistical significance.
Main Methods:
- The median p-value method was tested with nine common statistical tests, including t-tests, ANOVA, correlation, and regression.
- Empirical type I error rates were calculated for each test under varying levels of missing data.
- Performance was compared against the advertised statistical significance levels.
Main Results:
- The median p-value method resulted in inflated empirical type I error rates for all tested statistical analyses.
- The inflation of type I error rates increased proportionally with the prevalence of missing data.
- The method demonstrated a consistent overestimation of statistical significance.
Conclusions:
- The median p-value method is unreliable for assessing statistical significance when pooling test statistics from imputed datasets.
- Researchers should avoid using the median p-value method due to its tendency to produce false positives.
- Alternative, more robust methods are needed for pooling test statistics in multiply imputed data.
Related Concept Videos
Median
17.3K
Besides mean, the median is a widely used measure of central tendency. Typically, median is defined as the central or middle value of a data set, measured by arranging the data elements in an increasing or decreasing order. Since this middle value is not affected by the precise numerical values of the outliers or fluctuations, it is insensitive to them. Hence, in cases where a data set may have outliers or the extreme values are not known, the median is a better measure of the central tendency...
17.3K
Measures of Central Tendency
15.4K
The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean" and "average" is technically a center location. However, in practice among non-statisticians,...
15.4K
Sign Test for Median of Single Population
55
In general, the sign test serves as a nonparametric method to test hypotheses about the median of a single population when the data does not follow a known distribution. This simplicity makes it particularly useful for small sample sizes or when the assumptions of parametric tests cannot be met. The process begins with identifying a null hypothesis, typically stating that the population median equals a specific value. The alternative hypothesis could be that the median is either not equal to,...
55
Midrange
3.5K
A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to...
3.5K
Wilcoxon Signed-Ranks Test for Median of Single Population
64
The Wilcoxon signed-rank test for the median of a single population is a nonparametric test used to evaluate whether the median of a population differs from a specified value. Unlike parametric tests, it does not require data to follow a normal distribution, making it suitable for non-normal or small samples. The test begins by calculating the difference (d) between each observation and the hypothesized median. The absolute values of these differences are ranked in ascending order, with ties...
64
Review and Preview
6.8K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
6.8K

