Don't middle your MIDs: regression to the mean shrinks estimates of minimally important differences

Peter M Fayers1, Ron D Hays

  • 1Institute of Applied Health Sciences, University of Aberdeen, Aberdeen, UK, P.Fayers@abdn.ac.uk.

Insights

Estimating minimal important differences (MIDs) for patient-reported outcomes (PROs) using anchor-based clinical variables can be biased. Regression techniques for MID estimation are shown to be flawed, and alternative methods are recommended for accurate clinical interpretation.

Area of Science:

  • Clinical Epidemiology
  • Health Outcomes Research
  • Psychometrics

Background:

  • Patient-reported outcomes (PROs) are crucial for evaluating treatment effectiveness.
  • Minimal important differences (MIDs) quantify the smallest change in a PRO that patients perceive as beneficial.
  • Anchor-based methods are commonly used to estimate MIDs, linking PRO changes to clinical variables.

Purpose of the Study:

  • To critically evaluate the use of regression techniques for estimating MIDs based on clinical anchors.
  • To demonstrate the inherent bias in using regression for MID estimation.
  • To propose and advocate for alternative, less biased methods for calculating MIDs.

Main Methods:

  • The study theoretically analyzes the statistical properties of regression-based MID estimation.
  • It identifies sources of bias when using a clinical anchor to estimate PRO MIDs.
  • Alternative statistical approaches for anchor-based MID estimation are discussed.

Main Results:

  • Regression techniques applied to anchor-based MID estimation introduce significant bias.
  • This bias can lead to inaccurate estimations of clinically meaningful changes in PROs.
  • The findings highlight the limitations of current common practices.

Conclusions:

  • Regression-based methods for estimating MIDs from clinical anchors are inappropriate and should be avoided.
  • Accurate estimation of MIDs is essential for reliable interpretation of PRO data in clinical research and practice.
  • Further research should focus on validating and implementing alternative methods for robust MID determination.

Related Concept Videos

Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Midrange01:07

Midrange

A somewhat easy to compute quantitative estimate of a data set’s central tendency is its midrange, which is defined as the mean of the minimum and maximum values of an ordered data set.
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to outliers and...
Trimmed Mean01:10

Trimmed Mean

While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
Central Tendency: Analysis01:10

Central Tendency: Analysis

Measures of central tendency are tools used in biostatistics to identify the average or center of a dataset. They offer a single representative value for understanding and summarizing data distribution.
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
Standard Error of the Mean01:13

Standard Error of the Mean

The sampling variability of a statistic is defined as how much the statistic varies from one sample to another. The sampling variability of a statistic is typically measured by measuring its standard error.The standard error of the mean is an example of a standard error. It is a unique standard deviation known as the standard deviation of the sampling distribution of the mean. The standard error of the mean is a statistic that calculates how correctly a sample distribution represents a...
Testing a Claim about Mean: Known Population SD01:11

Testing a Claim about Mean: Known Population SD

A complete procedure of testing the hypothesis about a population mean is explained here.
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...