Related Experiment Video
Updated: Feb 8, 2026

Skeletal Muscle Gender Dimorphism from Proteomics
Published on: December 14, 2011
ABRF Proteome Informatics Research Group (iPRG) 2016 Study: Inferring Proteoforms from Bottom-up Proteomics Data.
Joon-Yong Lee1, Hyungwon Choi2, Christopher M Colangelo3
1Pacific Northwest National Laboratory, Richland, Washington 99352, USA.
The 2016 Proteome Informatics Research Group study evaluated proteoform inference and false discovery rate (FDR) estimation. This research provides a unique dataset for proteoform identification and method validation.
Area of Science:
- Biochemistry
- Bioinformatics
- Proteomics
Background:
- The Proteome Informatics Research Group (iPRG) conducted a study in 2016 to address challenges in proteoform inference and false discovery rate (FDR) estimation.
- Bottom-up proteomics data analysis presents complexities in accurately identifying and quantifying protein isoforms.
Purpose of the Study:
- To evaluate methods for proteoform inference and FDR estimation using bottom-up proteomics data.
- To assess participant performance in identifying proteins and estimating proteoform-level FDR.
- To test a new submission system for handling proteomics data analysis methods.
Main Methods:
- Generated triplicate Q Exactive Orbitrap liquid chromatography-tandem mass spectrometry datasets from four *Escherichia coli* samples.
- Spiked samples with equimolar mixtures of recombinant proteins to mimic homologous pairs.
- Provided participants with raw data and sequence files for proteoform identification and FDR estimation.
Main Results:
- The proteoform inference task was challenging, with only eight unique submissions received.
- No single method consistently outperformed others across all samples.
- Participants' submissions lacked complete executable R Markdown or IPython Notebooks, despite provided examples.
Conclusions:
- The study generated a unique "ground-truth" dataset for proteoform identification, now available to researchers.
- The developed virtual private server (VPS) and submission validator system are scalable for future studies.
- Enhanced promotion and participation are needed for future iPRG studies to maximize engagement and data generation.
Related Concept Videos
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Theory of Attribution I: Correspondent Inference Theory
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Data Reporting and Recording

