Related Experiment Video
Updated: Apr 28, 2026

10:05
Large-scale Top-down Proteomics Using Capillary Zone Electrophoresis Tandem Mass Spectrometry
Published on: October 24, 2018
8.9K
The C-score: a Bayesian framework to sharply improve proteoform scoring in high-throughput top down proteomics
Richard D LeDuc1, Ryan T Fellers, Bryan P Early
1National Center for Genome Analysis Support, Indiana University , 2709 E. 10th Street, Bloomington, Indiana 47408, United States.
Journal of Proteome Research
|June 13, 2014
Summary
We developed a new Bayesian scoring method, the C-score (Characterization Score), to improve protein identification and characterization in top-down proteomics. This method significantly enhances the accuracy of identifying complex protein forms (proteoforms).
Area of Science:
- Proteomics
- Biochemistry
- Computational Biology
Background:
- Automated analysis of top-down proteomics data requires enhanced scoring for accurate protein identification.
- Distinguishing highly related protein forms (proteoforms) remains a challenge in current proteomic analyses.
Purpose of the Study:
- To introduce the C-score (Characterization Score), a novel Bayesian approach for improved proteoform identification and characterization.
- To integrate expert knowledge into generative models for top-down proteomics data analysis.
Main Methods:
- Developed a Bayesian framework (C-score) incorporating protein properties and analytical system characteristics.
- Implemented intelligent weighting for site-specific modifications and accounted for fragmentation propensities and errors.
- Compared C-score performance against existing probability-based scoring systems (ProSightPC, ProSightPTM) using a curated dataset.
Main Results:
- The C-score framework demonstrated a marked improvement in scoring performance.
- Achieved a significantly higher area under the curve (AUC) of 0.99 compared to 0.78 for existing methods.
- Validated performance on a manually curated set of 295 human proteoforms.
Conclusions:
- The C-score offers a substantial advancement for proteoform identification and characterization in top-down proteomics.
- The Bayesian approach effectively leverages expert knowledge for more accurate data interpretation.
- This scoring system has the potential to improve the automated processing of complex proteomic datasets.

