Left-Censored Missing Value Imputation Approach for MS-Based Proteomics Data with GSimp

Runmin Wei1, Jingye Wang2

  • 1The University of Texas MD Anderson Cancer Center, Department of Genetics, Houston, TX, USA. rwei2@mdanderson.org.

Insights

Mass spectrometry (MS)-based omics studies often have missing values due to detection limits, which can bias results. GSimp, a new Gibbs sampler method, effectively imputes these missing values in MS-proteomics data.

Area of Science:

  • Biochemistry
  • Proteomics
  • Bioinformatics

Background:

  • Missing values are common in mass spectrometry (MS)-based omics, particularly proteomics.
  • Values below the limit of detection or quantification (LOD/LOQ) are often missing not at random (MNAR).
  • MNAR data can cause biased statistical analysis and hinder downstream applications.

Purpose of the Study:

  • To introduce GSimp, a novel imputation method for MS-proteomics data.
  • To address the challenge of missing not at random (MNAR) values in omics studies.
  • To provide a tool for accurate statistical estimation in MS-based omics.

Main Methods:

  • Development of GSimp, a Gibbs sampler-based imputation approach.
  • Focus on imputing left-censored missing values specific to MS-proteomics.
  • Detailed explanation of MNAR principles and GSimp's application.

Main Results:

  • GSimp effectively handles left-censored missing values in MS-proteomics datasets.
  • The method is designed to mitigate bias caused by MNAR data.
  • Facilitates more reliable downstream analyses in omics research.

Conclusions:

  • GSimp offers a robust solution for imputing MNAR values in MS-proteomics.
  • Accurate imputation is crucial for unbiased statistical inference in omics.
  • This approach supports the advancement of MS-based omics studies.

Related Concept Videos