The Protein Data Bank archive as an open data resource
Helen M Berman1, Gerard J Kleywegt, Haruki Nakamura
1RCSB PDB, Department of Chemistry and Chemical Biology and Center for Integrative Proteomics Research, Rutgers, The State University of New Jersey, Piscataway, NJ, 08854, USA, berman@rcsb.rutgers.edu.
Journal of Computer-Aided Molecular Design
|July 27, 2014
Summary
The Protein Data Bank, a vital biological data resource, has evolved over 40 years. Its enduring success stems from the interplay of science, technology, and community engagement.
Area of Science:
- Structural Biology
- Bioinformatics
- Open Science
Background:
- Established in 1971, the Protein Data Bank (PDB) is a key archive for biological macromolecular structure data.
- The PDB recently marked its 40th anniversary, highlighting its long-standing role in scientific research.
Observation:
- Analysis reveals intricate connections between scientific advancements, technological innovations, and community involvement.
- These interrelationships have been crucial for the PDB's sustained growth and relevance.
Findings:
- The PDB has become one of the most established and frequently utilized open-access data resources in the field of biology.
- Its evolution demonstrates a successful model for managing and disseminating scientific data.
Implications:
- Understanding the PDB's development offers insights into the sustainability of open-access scientific data repositories.
- This model can inform future strategies for biological data sharing and resource management.
Related Concept Videos
Protein Families
13.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
13.3K
Protein Families
3.5K
3.5K
Protein Networks
1.8K
1.8K
Protein Networks
3.7K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.7K
Conservation of Protein Domains Over Different Proteins
11.7K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.7K
Protein and Protein Structures
14.6K
14.6K


