Related Experiment Video
Updated: Aug 6, 2026

A Mass Spectrometry-Based Proteomics Approach for Global and High-Confidence Protein R-Methylation Analysis
Published on: April 28, 2022
SoftHybrid: A Hybrid Imputation Algorithm Optimized for Single-Cell Proteomics Data
Yixin Shi1,2, Simon Davis1, Philip D Charles1,3
1Target Discovery Institute, Centre for Medicines Discovery, Nuffield Department of Medicine, University of Oxford, Roosevelt Drive, OxfordOX3 7FZ, U.K.
Abstract:
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive missing-not-at-random (MNAR) sparsity. Existing imputation methods typically target either missing-at-random (MAR) or MNAR mechanisms, resulting in a trade-off between replicate consistency and preservation of biological variation, and are largely designed for bulk data. Here, we introduce SoftHybrid, a data-driven imputation framework that jointly models missingness and protein abundance to estimate the probability of MNAR, enabling continuous weighting between MAR- and MNAR-oriented strategies. SoftHybrid requires no external priors (cell type labels, group annotations, predefined missingness assumptions, etc.), enabling fully unsupervised applications. Across ground truth benchmarks and real single-cell proteomics data sets, SoftHybrid outperforms existing methods at low input and matches or exceeds their performance at the minibulk level. By preserving the proteomic structure and abundance accuracy, it enhances the recovery of biologically meaningful signals. SoftHybrid is implemented as an R package and is freely available at GitHub.

