Boosting the Speed and Accuracy of Protein Quantification Algorithms in Mass Spectrometry-Based Proteomics
Thang V Pham1,2,3, Chau T M Tran3, Alex A Henneman1,2
1Amsterdam UMC, location Vrije Universiteit Amsterdam, Department of Medical Oncology, OncoProteomics Laboratory, De Boelelaan 1117, 1081 HV Amsterdam, The Netherlands.
None:
Protein quantification is a crucial data processing step that combines quantitative values at the peptide or fragment level into protein levels in mass spectrometry-based proteomics. However, many of the current algorithms, including the state-of-the-art method MaxLFQ, do not scale well with the increasing number of samples, because of the limited system memory and algorithmic complexities. Here we introduce the iq format, a novel data structure designed to support very large data sets. We optimize existing quantification methods for both speed and memory usage. In particular, the new algorithms maxlfq-bit and rlm-cd significantly improve the base methods, MaxLFQ and the robust linear model, respectively, achieving orders of magnitude speed improvements for a large number of samples. The experimental result shows that the MaxLFQ algorithm achieves the highest accuracy, despite its comparatively higher computational cost. We also introduce a generic algorithm to boost the quantification accuracy of all methods by reducing the effect of noisy ion intensity traces. The experimental results show that the weighting approach improves the performance of all tested methods on a spike-in data set and a mixed species data set. The software implementation is publicly available in the R package iq from version 2.
Related Concept Videos
Tandem Mass Spectrometry
Mass Spectrometry: Overview
Mass Spectrometry of Amines
Mass Spectrometry: Isotope Effect
Improving Translational Accuracy
Uncertainty in Measurement: Accuracy and Precision


