Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Genome-wide association studies of missing metabolite measures from two population-based studies.

Genome biology·2026
Same author

Pathophysiology-based subtypes of individuals at high risk of type 2 diabetes: multi-omics profiles and lifestyle differences across subtypes.

Metabolism: clinical and experimental·2026
Same author

Dynamics of infection, vaccination and excess mortality during the COVID-19 pandemic among older individuals-a nationwide analysis.

European journal of epidemiology·2026
Same author

Abdominal Fat Measures Are Associated With Sex-Specific Prothrombotic Changes in Middle-Aged Adults.

Arteriosclerosis, thrombosis, and vascular biology·2026
Same author

Distinct plasma protein profiles after long-term remission of Cushing's disease.

The Journal of clinical endocrinology and metabolism·2026
Same author

Insomnia as a risk factor for the development of depression and anxiety in primary care: a matched population-based cohort study.

Family practice·2026

Related Experiment Video

Updated: Nov 28, 2025

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
05:35

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome

Published on: September 20, 2022

4.0K

A Workflow for Missing Values Imputation of Untargeted Metabolomics Data.

Tariq Faquih1, Maarten van Smeden2, Jiao Luo1

  • 1Department of Clinical Epidemiology, Leiden University Medical Center, Postal Zone C7-P, PO Box 9600, 2300 RC Leiden, The Netherlands.

Metabolites
|December 1, 2020
PubMed
Summary

Missing data in metabolomics studies can bias results. This study offers an R script for imputing missing values, showing Multivariate Imputation by Chained Equations (MICE) performs best with larger sample sizes.

Keywords:
imputationk-nearest neighborsmetabolonmultiple imputation using chained equationssimulationuntargeted metabolomicsworkflow

More Related Videos

Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain
07:10

Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain

Published on: March 13, 2020

10.4K
A Strategy for Sensitive, Large Scale Quantitative Metabolomics
14:18

A Strategy for Sensitive, Large Scale Quantitative Metabolomics

Published on: May 27, 2014

21.4K

Related Experiment Videos

Last Updated: Nov 28, 2025

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
05:35

An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome

Published on: September 20, 2022

4.0K
Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain
07:10

Untargeted Liquid Chromatography-Mass Spectrometry-Based Metabolomics Analysis of Wheat Grain

Published on: March 13, 2020

10.4K
A Strategy for Sensitive, Large Scale Quantitative Metabolomics
14:18

A Strategy for Sensitive, Large Scale Quantitative Metabolomics

Published on: May 27, 2014

21.4K

Area of Science:

  • Metabolomics
  • Bioinformatics
  • Statistical Genetics

Background:

  • Metabolomics studies are expanding due to advanced platforms.
  • Missing data in large metabolite panels can lead to biased results if not handled properly.
  • Accurate imputation of missing values is crucial for reliable metabolomics research.

Purpose of the Study:

  • To develop and evaluate a user-friendly R script for imputing missing values in untargeted metabolomics data.
  • To assess the performance of Multivariate Imputation by Chained Equations (MICE) and k-nearest neighbors (kNN) imputation methods.
  • To investigate the impact of sample size, missing data percentage, and correlation structure on imputation accuracy.

Main Methods:

  • Developed a publicly available R script for metabolite imputation.
  • Evaluated MICE and kNN algorithms using simulated missing data.
  • Simulations were based on real metabolomics data from the Netherlands Epidemiology of Obesity (NEO) study (n=599), varying sample sizes, missing percentages, and missing mechanisms.

Main Results:

  • For MICE, larger sample size significantly reduced bias and error in imputation.
  • For kNN, imputation accuracy was primarily influenced by the correlation between metabolites.
  • MICE demonstrated superior performance, especially with larger datasets (n > 50).

Conclusions:

  • An effective imputation workflow for untargeted metabolomics data is presented via a public R script.
  • Simulation results offer critical insights into factors affecting imputation accuracy, guiding method selection.
  • The study highlights the importance of sample size for MICE and metabolite correlation for kNN in metabolomics data analysis.