Related Experiment Videos
Simple sequences are rare in the Protein Data Bank.
Melanie A Huntley1, G Brian Golding
1Department of Biology, McMaster University, Hamilton, Ontario, Canada.
Proteins
|May 16, 2002
Summary
Highly repetitive protein sequences are common but rarely structurally characterized. This study suggests these simple sequences may form disordered protein structures, hindering their inclusion in structural databases.
Area of Science:
- Protein bioinformatics
- Structural biology
- Genomics
Background:
- Simple sequence repeats (SSRs) are abundant in sequenced proteomes.
- Highly repetitive SSRs are prevalent in eukaryotes but scarce in prokaryotes.
Purpose of the Study:
- To investigate the representation of low complexity, highly repetitive protein sequences in the Protein Data Bank (PDB).
- To explore the structural characteristics of these repetitive sequences within known protein structures.
Main Methods:
- Comparative analysis of eukaryotic proteins in the PDB versus simulated databases derived from the National Center for Biotechnology Information (NCBI).
- Detailed examination of structural data for PDB entries containing highly repetitive simple sequences.
Main Results:
- The PDB shows a significant underrepresentation of low complexity, highly repetitive protein sequences compared to simulated NCBI databases.
- For the few identified repetitive sequences in the PDB, tertiary structure information is often missing for these regions.
Conclusions:
- Highly repetitive simple sequences are deficient in structural databases like the PDB.
- The lack of structural characterization suggests these sequences may form intrinsically disordered protein (IDP) regions, complicating structural determination.