Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

7.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
7.1K
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

16.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
16.7K
Outliers and Influential Points01:08

Outliers and Influential Points

6.5K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.5K
What Are Outliers?01:12

What Are Outliers?

5.4K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
5.4K
Censoring Survival Data01:09

Censoring Survival Data

609
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
609
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

4.2K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Explainable Machine Learning Analysis of Perioperative Factors Associated with Clinically Significant Emergence Agitation After Pediatric Ophthalmic Surgery.

Medicina (Kaunas, Lithuania)·2026
Same author

Machine Learning-Based Prediction of Early Patient-Controlled Analgesia Discontinuation After Total Knee Arthroplasty: A Retrospective Cohort Study.

Journal of clinical medicine·2026
Same author

Prediction of Surgical Intervention in Acute Knee Trauma: A Focus on Threshold-Specific Performance and Clinical Decision Utility.

Diagnostics (Basel, Switzerland)·2026
Same author

Opioid-Sparing Effect of Celiac Plexus Neurolysis in Palliative Care Patients with Upper Abdominal Cancer: A Single-Center Retrospective Case Series of Thirteen Patients.

Journal of hospice and palliative care·2026
Same author

Mode of anaesthesia and persistent postoperative opioid use: A reply.

Anaesthesia·2026
Same author

Disease Stage-Dependent Association Between Nephrotic-Range Proteinuria and Severe Acute Kidney Injury in Patients with Liver Cirrhosis.

Journal of clinical medicine·2026

Related Experiment Video

Updated: Feb 24, 2026

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
08:51

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts

Published on: September 20, 2024

2.2K

Statistical data preparation: management of missing values and outliers.

Sang Kyu Kwak1, Jong Hae Kim2

  • 1Department of Medical Statistics, School of Medicine, Catholic University of Daegu, Daegu, Korea.

Korean Journal of Anesthesiology
|August 11, 2017
PubMed
Summary

Handling missing data and outliers is crucial for reliable research. This review covers identifying and managing these common data issues to ensure accurate statistical analysis and trustworthy results.

Keywords:
BiasData collectionData interpretationStatistics

More Related Videos

Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain
07:10

Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain

Published on: March 13, 2020

10.9K
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.4K

Related Experiment Videos

Last Updated: Feb 24, 2026

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
08:51

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts

Published on: September 20, 2024

2.2K
Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain
07:10

Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain

Published on: March 13, 2020

10.9K
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.4K

Area of Science:

  • Data Science
  • Statistics
  • Research Methodology

Background:

  • Missing values and outliers are common challenges in data collection.
  • These data anomalies can compromise statistical power, introduce bias, and reduce data efficiency.
  • Improper handling of missing values and outliers significantly impacts study reliability and results.

Purpose of the Study:

  • To review different types of missing values.
  • To discuss methods for identifying outliers in datasets.
  • To provide strategies for effectively dealing with missing values and outliers.

Main Methods:

  • Literature review of data imputation and outlier detection techniques.
  • Exploration of statistical methods for identifying anomalies.
  • Synthesis of approaches for handling missing data and outliers.

Main Results:

  • Categorization of missing data into different types.
  • Overview of various outlier detection algorithms and statistical tests.
  • Discussion of imputation methods and outlier treatment strategies.

Conclusions:

  • Effective management of missing values and outliers is essential for robust data analysis.
  • Choosing appropriate methods depends on the data characteristics and research goals.
  • Proper data preprocessing enhances the validity and reliability of research findings.