Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

What Are Outliers?01:12

What Are Outliers?

Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the vertical...
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Modified Boxplots00:57

Modified Boxplots

A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
Trimmed Mean01:10

Trimmed Mean

While measuring the mean of a data set, care needs to be taken when associating the mean to its central tendency. The same goes for the arithmetic mean, the geometric mean, or the harmonic mean. This is because the presence of a single outlier data value can significantly affect the mean. That is, the mean is sensitive to fluctuations in the data set.
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Vascular geometry and oxygen diffusion in the vicinity of artery-vein pairs in the kidney.

American journal of physiology. Renal physiology·2014
Same author

Letter to the editor. Second thoughts on the sidedness of P.

Clinical and experimental pharmacology & physiology·2013
Same author

Telemetry-based oxygen sensor for continuous monitoring of kidney oxygenation in conscious rats.

American journal of physiology. Renal physiology·2013
Same author

Should we use one-sided or two-sided P values in tests of significance?

Clinical and experimental pharmacology & physiology·2013
Same author

Analysing 2 × 2 contingency tables: which test is best?

Clinical and experimental pharmacology & physiology·2013
Same author

A primer for biomedical scientists on how to execute model II linear regression analysis.

Clinical and experimental pharmacology & physiology·2011

Related Experiment Videos

Outlying observations and missing values: how should they be handled?

John Ludbrook1

  • 1Department of Surgery, The University of Melbourne, Melbourne, Victoria, Australia. ludbrook@bigpond.net.au

Clinical and Experimental Pharmacology & Physiology
|January 25, 2008
PubMed
Summary

Handling outlying observations and missing values in research depends on group size. Graphical methods like scatterplots and box-and-whisker plots aid detection. Strategies vary for small versus large groups to ensure data integrity.

Related Experiment Videos

Area of Science:

  • Biostatistics
  • Clinical Research Methodology
  • Data Analysis

Background:

  • Outlying observations and missing values present challenges in data analysis.
  • The appropriate handling of these data issues is contingent upon experimental group sizes.
  • Small groups (3-44 observations) and large groups (100s-1000s observations) require distinct strategies.

Purpose of the Study:

  • To delineate effective strategies for managing outlying observations and missing values in research.
  • To provide guidance on data handling techniques tailored to different experimental group sizes.
  • To ensure the integrity and validity of statistical analyses in clinical and experimental studies.

Main Methods:

  • Detection of outlying observations using graphical methods such as scatterplots (mean+/-2 SD) and box-and-whisker plots.
  • Evaluation of strategies for handling confirmed outlying observations, including deletion (treated as missing values) or replacement if an independent explanation exists.
  • Exploration of methods for unexplained extreme values, such as data transformation, permutation tests, and robust statistical methods (median, trimmed mean, etc.).
  • Review of techniques for addressing missing values, including ignoring, manual imputation (small datasets), and computerized imputation (large datasets, e.g., regression, EM algorithms).

Main Results:

  • Graphical methods are recommended for detecting outlying observations, though interpretation requires judgment.
  • Deletion of outlying observations is permissible only if an independent explanation is identified.
  • Unexplained extreme values necessitate alternative analytical approaches, including data transformation or robust statistical methods.
  • Missing value imputation techniques are crucial for large datasets, but require careful assessment for potential bias.

Conclusions:

  • The optimal approach for managing outlying observations and missing values is critically dependent on the size of the experimental groups.
  • Robust statistical methods and appropriate imputation techniques are essential for maintaining analytical rigor, particularly in large-scale studies.
  • It is imperative to test for non-random missingness to prevent bias in research findings.