Related Experiment Video
Updated: Feb 20, 2026

10:37
Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
Published on: November 15, 2017
12.8K
Practical Impact of Imputation and Batch-Effect Correction for Proteomics/Peptidomics Differential-Abundance Analysis
Charis Gonidaki1,2, Agnieszka Latosinska3, Antonia Vlahou1
1Center of Systems Biology, Biomedical Research Foundation of the Academy of Athens (BRFAA), Athens, Greece.
Proteomics
|February 18, 2026
Summary
Choosing the right data preprocessing methods is crucial in clinical proteomics. Batch correction, especially without considering disease status, can remove significant biological signals, impacting biomarker discovery in chronic kidney disease research.
Area of Science:
- Proteomics
- Biomarker Discovery
- Clinical Chemistry
Background:
- Mass spectrometry (MS)-based proteomics is valuable for biomarker discovery but faces challenges like missing values and batch effects.
- Existing imputation and batch-correction methods' impact on large clinical proteomics datasets is not fully understood.
- Chronic kidney disease (CKD) biomarker research requires robust data analysis pipelines.
Purpose of the Study:
- To evaluate the practical impact and interaction of common imputation and batch-correction methods on differential abundance analysis.
- To assess these methods' effects on identifying disease-associated peptides (DAPs) in a large-scale CKD urine peptidomics dataset.
- To understand how preprocessing choices influence biological signal preservation in clinical proteomics.
Main Methods:
- Analyzed a CE-MS urine peptidomics dataset with 1,050 samples across 13 batches from CKD patients and controls.
- Tested three imputation methods (Gaussian, ½ LOD, KNN) combined with three batch-correction approaches (ComBat, ComBat with disease covariate, MNN).
- Assessed downstream effects by validating identified DAPs between discovery and validation sets.
Main Results:
- Imputation method choice had minimal impact on the final list of DAPs.
- Batch-effect correction significantly altered results; MNN and unadjusted ComBat removed substantial proportions of DAPs (~50% and >90%).
- Including disease status in the ComBat model largely preserved the biological signal compared to unadjusted methods.
Conclusions:
- Preprocessing choices, particularly batch-effect correction strategies, critically influence downstream results in clinical proteomics.
- Imputation and batch-effect correction methods interact, jointly affecting the identification of DAPs.
- Applying batch-effect correction requires caution, as it can lead to the loss of meaningful biological signal, emphasizing the need for careful model selection.

