Related Experiment Video
Updated: Jul 9, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Annotating Macromolecular Complexes in the Protein Data Bank: Improving the FAIRness of Structure Data
Sri Devan Appasamy1, John Berrisford2, Romana Gaborova3
1Protein Data Bank in Europe, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge, CB10 1SD, UK. sria@ebi.ac.uk.
Standardizing macromolecular complex names in the Protein Data Bank (PDB) improves data accessibility. This method uses external resources to assign unique identifiers, enhancing research and understanding of biological assemblies.
Area of Science:
- Structural Biology
- Bioinformatics
- Molecular Biology
Background:
- Macromolecular complexes are vital for cellular functions.
- Atomic-level understanding of these complexes is crucial for biological research.
- The Protein Data Bank (PDB) stores structural data but lacks standardized assembly naming.
Purpose of the Study:
- To develop a method for standardizing the naming of macromolecular assemblies in the PDB.
- To improve the identification and contextualization of biological assemblies.
- To enhance the PDB's data quality and FAIR (Findable, Accessible, Interoperable, Reusable) attributes.
Main Methods:
- Leveraging external databases like Complex Portal, UniProt, and Gene Ontology.
- Developing a systematic approach to assign standard names and persistent identifiers to PDB assemblies.
- Applying the method to a large dataset of unique assemblies within the PDB.
Main Results:
- Successfully assigned standard names to over 90% of unique PDB assemblies.
- Provided persistent identifiers for each standardized assembly.
- Demonstrated significant improvement in data standardization and contextualization.
Conclusions:
- The developed method effectively standardizes macromolecular assembly data in the PDB.
- Enhanced data standardization improves the PDB's utility for research and education.
- Improved FAIR attributes of PDB data facilitate basic and translational research.
More Related Videos
14:55Atomic Scale Structural Studies of Macromolecular Assemblies by Solid-state Nuclear Magnetic Resonance Spectroscopy
Published on: September 17, 2017
07:33Analyzing Protein Architectures and Protein-Ligand Complexes by Integrative Structural Mass Spectrometry
Published on: October 15, 2018
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Globular and Fibrous Proteins
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...