Related Experiment Video
Updated: Feb 25, 2026

Enhanced Sample Multiplexing of Tissues Using Combined Precursor Isotopic Labeling and Isobaric Tagging cPILOT
Published on: May 1, 2017
Statistical Models for the Analysis of Isobaric Tags Multiplexed Quantitative Proteomics.
Gina D'Angelo1, Raghothama Chaerkady1, Wen Yu1
1Statistical Sciences, ‡Antibody Discovery and Protein Engineering, Protein Sciences, §Research Bioinformatics, ∥Clinical Biomarkers and Computational Biology, and ⊥Clinical Pharmacology, Pharmacometrics, and DMPK, MedImmune , Gaithersburg, Maryland 20878, United States.
This study compares statistical models for analyzing complex proteomic data from mass spectrometry. Median sweeping normalization with mixed models offers the best error control for identifying protein biomarkers.
Area of Science:
- Proteomics
- Biomarker Discovery
- Statistical Bioinformatics
Background:
- Mass spectrometry is crucial for identifying protein biomarkers for drug development.
- Proteomic data is complex, hierarchical, and often involves small sample sizes.
- Generalized linear models (GLM) are common but don't handle all data complexities like repeated measures or variance heterogeneity.
Purpose of the Study:
- To compare the performance of three statistical models: generalized linear models (GLM), linear models for microarray data (LIMMA), and mixed models.
- To evaluate these models under two normalization methods: quantile normalization and median sweeping.
- To determine the optimal statistical approach for analyzing tandem mass tag (TMT) labeled proteomic data.
Main Methods:
- Comparison of GLM, LIMMA, and mixed models.
- Application of quantile normalization and median sweeping normalization techniques.
- Evaluation using spiked-in data, a systemic lupus erythematosus (SLE) dataset, and simulated TMT data.
Main Results:
- Median sweeping was identified as a preferred normalization approach.
- GLM results were found to be a subset of mixed model findings when using median sweeping.
- Mixed models demonstrated the best type I error control with median sweeping normalization.
Conclusions:
- The mixed model combined with median sweeping normalization provides superior control over type I errors in proteomic data analysis.
- LIMMA exhibited better overall statistical properties irrespective of the normalization method employed.
- Choosing the appropriate statistical model and normalization strategy is critical for accurate protein biomarker identification.

