Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Prediction Intervals01:03

Prediction Intervals

2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.3K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Survival Tree01:19

Survival Tree

93
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
93
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.6K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

142
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
142
Improving Translational Accuracy02:07

Improving Translational Accuracy

11.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A disentangled transformer-based transfer learning framework to predict patient drug response from tumor single-cell transcriptomics.

Bioinformatics (Oxford, England)·2026
Same author

Pelvic MRI in Endometrial Cancer Staging: A Retrospective Evaluation of the Impact of FIGO 2023.

Radiology·2026
Same author

Blurring evidence with advocacy: a systematic review of policy recommendations for net zero.

npj environmental social sciences·2026
Same author

PhotIQA: A photoacoustic image data set with image quality ratings.

Scientific data·2026
Same author

Caregiver-Associated Physical Activity Patterns, Dietary Behaviors and Interventional Beliefs in Individuals with Down Syndrome: Insights from a Large European Survey.

Nutrients·2026
Same author

Understanding Obesity in Individuals with Down Syndrome: Caregiver Perceptions, Awareness, and Motivation.

Nutrients·2026

Related Experiment Video

Updated: Jul 14, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K

The impact of imputation quality on machine learning classifiers for datasets with missing values.

Tolou Shadbahr1, Michael Roberts2,3, Jan Stanczuk4

  • 1Research Program in Systems Oncology, Faculty of Medicine, University of Helsinki, Helsinki, Finland.

Communications Medicine
|October 6, 2023
PubMed
Summary

Classifying incomplete datasets requires careful data imputation. Poor imputation quality significantly degrades classifier performance, highlighting the need to prioritize accurate imputation methods for reliable machine learning outcomes.

More Related Videos

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K
Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
04:09

Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma

Published on: October 10, 2018

8.3K

Related Experiment Videos

Last Updated: Jul 14, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K
Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
04:09

Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma

Published on: October 10, 2018

8.3K

Area of Science:

  • Machine Learning
  • Data Science
  • Statistics

Background:

  • Classifying samples in incomplete datasets is a common machine learning challenge.
  • Real-world datasets often contain missing values requiring imputation before classification.
  • Current focus is on optimizing classifier performance after imputation.

Purpose of the Study:

  • To evaluate the impact of imputation methods and missingness rates on downstream classifier performance.
  • To compare existing imputation quality assessment methods with a novel approach using sliced Wasserstein distance.
  • To analyze the stability and interpretability of models trained on imputed data.

Main Methods:

  • Utilized three simulated and three real-world clinical datasets with varying missingness patterns.
  • Employed Analysis of Variance (ANOVA) to quantify the influence of missingness rate, imputation, and classifier choices.
  • Introduced and evaluated discrepancy scores based on sliced Wasserstein distance for imputation quality assessment.

Main Results:

  • Classifier performance is highly sensitive to the percentage of missingness in test data.
  • Common imputation quality measures often result in data distributions that poorly match the original.
  • Novel discrepancy scores demonstrate superior performance in matching underlying data distributions.
  • Interpretability of classifier models is compromised when trained on poorly imputed data.

Conclusions:

  • The quality of data imputation is critical for downstream classification tasks.
  • Poor imputation can lead to considerable negative effects on classifier performance and model interpretability.
  • Prioritizing robust imputation quality assessment is imperative for reliable machine learning applications.