PDB NextGen Archive: centralizing access to integrated annotations and enriched structural information by the
Preeti Choudhary1, Zukang Feng2, John Berrisford1
1Protein Data Bank in Europe, European Molecular Biology Laboratory, European Bioinformatics Institute Wellcome Genome Campus, Hinxton, Cambridgeshire, CB10 1SD, UK.
Summary
The Protein Data Bank (PDB) Next Generation Archive centralizes structural annotations, improving access to biomolecular data. This resource enhances research by providing integrated, up-to-date information from trusted sources.
Area of Science:
- Structural Biology
- Bioinformatics
- Data Archiving
Background:
- The Protein Data Bank (PDB) is a global repository for 3D biomolecular structures.
- Integrating external annotations into the PDB is challenging due to its archival nature and distributed data centers.
Purpose of the Study:
- To establish a centralized system for enriched structural annotations.
- To streamline access to updated information from PDB partners and external biodata resources.
Main Methods:
- Development of the PDB Next Generation (NextGen) Archive.
- Centralization of mappings between PDB structures, UniProt sequences, and domain annotations (Pfam, SCOP2, CATH).
- Inclusion of intra-molecular connectivity information.
Main Results:
- The NextGen Archive provides integrated data, including protein structures, UniProt sequences, and domain annotations.
- Substantial user engagement with over 3.5 million data file downloads since launch.
- Ensured researchers have access to accurate, up-to-date, and easily accessible structural annotations.
Conclusions:
- The PDB NextGen Archive successfully centralizes and streamlines access to enriched structural annotations.
- The archive enhances research by providing researchers with reliable and integrated biomolecular data.
- High user engagement indicates the value and utility of the NextGen Archive in the scientific community.
Related Concept Videos
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Next-generation Sequencing
88.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.7K
Nucleic Acid Structure
6.1K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
6.1K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Gene Families
8.8K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.8K


