Outlier Detection in Mendelian Randomization

Maximilian M Mandl1,2, Anne-Laure Boulesteix1,2, Stephen Burgess3,4

  • 1Institute for Medical Information Processing, Biometry, and Epidemiology, Faculty of Medicine, Ludwig-Maximilians-Universität, München, Germany.

PubMed

Insights

Mendelian randomization (MR) methods can over-identify genetic outliers due to pleiotropy. This study introduces a novel method to correct for overdispersion in heterogeneity statistics, improving the accuracy of causal inference from genetic data.

Area of Science:

  • Genetics
  • Epidemiology
  • Biostatistics

Background:

  • Mendelian randomization (MR) infers causal relationships using genetic variants as instrumental variables.
  • A core MR assumption is instrumental variable independence from outcomes, except via the exposure.
  • Pleiotropy, where variants affect outcomes through other pathways, violates this assumption and is common.

Purpose of the Study:

  • To address the overdispersion issue in heterogeneity statistics used to detect outlying genetic instruments in MR.
  • To develop a method for accurately identifying and removing pleiotropic instruments in Mendelian randomization analyses.

Main Methods:

  • Proposed a novel statistical method to correct for overdispersion in heterogeneity statistics.
  • Utilized an estimated inflation factor to identify and remove outlying genetic variants.
  • The method is applicable to both univariable and multivariable Mendelian randomization.

Main Results:

  • The new method effectively corrects for overdispersion in heterogeneity statistics.
  • Accurate removal of outlying instruments due to pleiotropy was achieved.
  • Improved reliability of causal effect estimates in Mendelian randomization.

Conclusions:

  • The developed method enhances the robustness of Mendelian randomization by accurately accounting for pleiotropic effects.
  • This approach improves the identification of valid genetic instruments, leading to more reliable causal inference.
  • The method is suitable for use with readily available summary-level genetic data.

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.1K
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.2K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K
What Are Outliers?01:12

What Are Outliers?

Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
4.2K
Modified Boxplots00:57

Modified Boxplots

A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
10.1K
Chi-square Analysis02:46

Chi-square Analysis

The chi-square test is a statistical hypothesis test. It is used to check whether there is a significant difference between an expected value and an observed value. In the context of genetics, it enables us to either accept or reject a hypothesis, based on how much the observed values deviate from the expected values.
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
38.7K