PaxDb v6.0: reprocessed, LLM-selected, curated protein abundance data across organisms
Qingyao Huang1,2, Damian Szklarczyk1,2, John Oehninger2
1Swiss Institute of Bioinformatics, Winterthurerstrasse 190, 8057 Zurich, Switzerland.
Nucleic Acids Research
|November 3, 2025
Summary
The PaxDb v6.0 database offers a standardized protein abundance reference for 392 species, improving reproducibility and data reuse from mass spectrometry (MS) experiments. This resource enhances biological insights by integrating reprocessed public MS data with consistent metadata.
Area of Science:
- Proteomics
- Bioinformatics
- Systems Biology
Background:
- Mass spectrometry (MS) data reuse is hindered by inconsistent processing and metadata, limiting biological discovery.
- Standardized, high-coverage reference resources are crucial for reproducibility and cross-study integration in proteomics.
- PaxDb provides organism- and tissue-level protein abundance data for the healthy, wild-type state.
Purpose of the Study:
- To enhance the PaxDb database with updated and expanded protein abundance data.
- To develop a standardized, automated pipeline for reprocessing public mass spectrometry data.
- To improve the accessibility and utility of public proteomics data for biological research.
Main Methods:
- Integrated 1639 datasets from 392 species into PaxDb v6.0.
- Developed an end-to-end MS data processing pipeline using the FragPipe framework for consistent re-analysis.
- Implemented standardized metadata integration, orthology mappings, and quality scoring.
- Utilized large-language model ensemble classifiers for semi-automated curation of ProteomeXchange projects.
- Created a user-facing tool for peptide-level abundance calculation and dataset comparison.
Main Results:
- PaxDb v6.0 significantly expands coverage across all kingdoms of life, nearly doubling previous dataset integration.
- A novel automated pipeline ensures consistent reprocessing of raw MS data, enhancing data quality and reliability.
- The updated database includes standardized metadata, orthology information, and quality scores for integrated datasets.
- New tools facilitate direct comparison of user data with PaxDb reference proteomes.
Conclusions:
- PaxDb v6.0 represents a significant advancement in creating a standardized, high-coverage protein abundance reference resource.
- The automated reprocessing pipeline and enhanced data integration overcome key limitations in public MS data reuse.
- This updated resource facilitates new biological insights by improving reproducibility and enabling cross-study data integration.
Related Concept Videos
Proteomics
9.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
9.3K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
Protein Networks
2.8K
2.8K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K


