Related Experiment Video
Updated: Aug 10, 2026

07:38
Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019
Challenges in deriving high-confidence protein identifications from data gathered by a HUPO plasma proteome
David J States1, Gilbert S Omenn, Thomas W Blackwell
1University of Michigan, 100 Washtenaw Rd., Palmer Commons 2035B, Ann Arbor, Michigan 48109, USA.
Nature Biotechnology
|March 10, 2006
Summary
The Human Proteome Organization identified 889 high-confidence human serum and plasma proteins using integrated mass spectrometry data. This comprehensive proteome characterization advances understanding of human proteins and novel gene sequences.
Area of Science:
- Proteomics
- Human Biology
- Biochemistry
Background:
- The Human Proteome Organization (HUPO) conducted a large-scale collaborative study.
- Characterizing the human serum and plasma proteomes is crucial for understanding human biology.
Purpose of the Study:
- To characterize the human serum and plasma proteomes through a collaborative, multi-laboratory effort.
- To establish a high-confidence protein dataset using integrated proteomic data analysis.
Main Methods:
- Utilized liquid chromatography-tandem mass spectrometry (LC-MS/MS) across 18 laboratories.
- Integrated and statistically analyzed proteomic data sets against the International Protein Index database.
- Employed rigorous statistical approaches, including accounting for coding region length and multiple hypothesis testing.
Main Results:
- Initially identified 9,504 proteins with one or more peptides and 3,020 with two or more peptides.
- Reduced the dataset to 889 proteins identified with at least 95% confidence using advanced statistical methods.
- Demonstrated the value of integrated analysis for accurate proteome representation.
Conclusions:
- A high-confidence set of 889 human serum and plasma proteins was established.
- Integrated proteomic data analysis provides an accurate representation of the proteome.
- The dataset facilitates high-confidence identification of protein matches to novel exons and non-annotated gene sequences.