Related Experiment Videos
Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories
Lennart Martens1, Alexey I Nesvizhskii, Henning Hermjakob
1Department of Biochemistry, Faculty of Medicine and Health Sciences, Ghent University, Ghent, Belgium. lennart.martens@UGent.be
Proteomics
|July 26, 2005
Summary
The Human Plasma Proteome Project (PPP) pilot phase is complete, highlighting the need for centralized proteomics data storage. Researchers discuss storing raw mass spectrometry data versus processed peak lists for better data dissemination and future research.
Area of Science:
- Proteomics
- Bioinformatics
- Data Science
Background:
- The Human Plasma Proteome Project (PPP) pilot phase has concluded, marking a significant milestone in proteomics research.
- The large volume of data generated necessitates a robust, centralized data dissemination mechanism.
- Existing proteomics data repositories and infrastructure development are crucial for managing this data.
Purpose of the Study:
- To address the debate on storing raw mass spectrometry (MS) data versus processed peak lists within the PPP.
- To detail the advantages and disadvantages of centralized storage for both raw and processed proteomics data.
- To provide recommendations for immediate and future MS data storage in public repositories.
Main Methods:
- Analysis of data management strategies employed during the PPP pilot phase.
- Evaluation of the merits and caveats of storing raw binary MS data.
- Assessment of the benefits and drawbacks of storing processed peak lists.
Main Results:
- The PPP pilot phase generated a substantial amount of proteomics data, underscoring data management challenges.
- A centralized data gathering infrastructure was developed at the University of Michigan, Ann Arbor.
- The European Bioinformatics Institute established a protein identifications database as a general proteomics data repository.
Conclusions:
- The choice between storing raw data and peak lists has significant implications for proteomics data accessibility and future research.
- Centralized storage of proteomics data, whether raw or processed, is essential for the scientific community.
- Recommendations are proposed for optimizing MS data storage in public repositories for long-term scientific benefit.