Related Experiment Videos
Identification of differentially expressed peptides in high-throughput proteomics data
Michiel P van Ooijen1, Victor L Jong2, Marinus J C Eijkemans3
1Department of Viroscience, Erasmus MC, CA Rotterdam, Netherlands.
Briefings in Bioinformatics
|April 4, 2017
Summary
High-throughput proteomics requires robust statistical methods for peptide-level analysis. The empirical Bayes method (limma) offers the highest sensitivity for differential expression analysis, outperforming t-tests and ANOVA.
Area of Science:
- Proteomics
- Bioinformatics
- Statistical Analysis
Background:
- High-throughput proteomics generates vast datasets, challenging quantitative analysis.
- Peptide-level analysis offers deeper insights into sub-protein variations like splice variants and post-translational modifications.
- Current statistical methods (t-test, ANOVA) often rely on data imputation, with limited evaluation due to lack of gold standards.
Purpose of the Study:
- To evaluate statistical methods for label-free, peptide-based differential proteomics data.
- To assess the impact of data imputation on statistical analysis performance.
- To determine optimal biological replicate numbers for reliable high-throughput data analysis.
Main Methods:
- Comparison of four statistical analysis methods on experimental and resampled proteomics data.
- Evaluation of data imputation techniques across varying numbers of biological replicates.
- Performance assessment based on sensitivity and false discovery rates.
Main Results:
- Three to four biological replicates are essential for confident identification of significant changes.
- Data imputation can increase sensitivity but significantly elevates the false discovery rate.
- The empirical Bayes method (limma) demonstrated superior sensitivity for peptide-level differential expression analysis.
Conclusions:
- Recommends the empirical Bayes method (limma) for peptide-level differential expression analysis in high-throughput proteomics.
- Highlights the importance of sufficient biological replicates for robust statistical findings.
- Warns of the increased false discovery rate associated with data imputation.