Related Experiment Video
Updated: Oct 20, 2025

Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
Published on: November 15, 2017
Multiple Imputation Approaches Applied to the Missing Value Problem in Bottom-Up Proteomics
Miranda L Gardner1,2, Michael A Freitas1,2
1Ohio State Biochemistry Program, Chemistry and Biochemistry, The Ohio State University, Columbus, OH 43210, USA.
Abstract:
Analysis of differential abundance in proteomics data sets requires careful application of missing value imputation. Missing abundance values widely vary when performing comparisons across different sample treatments. For example, one would expect a consistent rate of "missing at random" (MAR) across batches of samples and varying rates of "missing not at random" (MNAR) depending on the inherent difference in sample treatments within the study. The missing value imputation strategy must thus be selected that best accounts for both MAR and MNAR simultaneously. Several important issues must be considered when deciding the appropriate missing value imputation strategy: (1) when it is appropriate to impute data; (2) how to choose a method that reflects the combinatorial manner of MAR and MNAR that occurs in an experiment. This paper provides an evaluation of missing value imputation strategies used in proteomics and presents a case for the use of hybrid left-censored missing value imputation approaches that can handle the MNAR problem common to proteomics data.
Insights
Accurate proteomics analysis demands effective missing value imputation. This study advocates for hybrid left-censored methods to address both missing at random (MAR) and missing not at random (MNAR) data, crucial for reliable differential abundance findings.
Area of Science:
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- Differential abundance analysis in proteomics is challenged by missing data.
- Missing values can arise from missing at random (MAR) or missing not at random (MNAR) mechanisms.
- Existing imputation strategies may not adequately address both MAR and MNAR simultaneously.
Purpose of the Study:
- To evaluate existing missing value imputation strategies for proteomics data.
- To identify imputation methods that can effectively handle both MAR and MNAR data.
- To present a case for hybrid left-censored imputation approaches.
Main Methods:
- Evaluation of various missing value imputation strategies.
- Focus on methods addressing the combinatorial nature of MAR and MNAR.
- Analysis of proteomics datasets with varying missing value patterns.
Main Results:
- Missing value patterns in proteomics data are complex, involving both MAR and MNAR.
- The rate of MNAR can be influenced by experimental sample treatments.
- Hybrid left-censored imputation demonstrates potential for handling MNAR.
Conclusions:
- Appropriate selection of missing value imputation is critical for reliable proteomics analysis.
- Hybrid left-censored imputation methods offer a promising solution for MNAR challenges.
- Further adoption of these methods can improve the accuracy of differential abundance studies.

