Related Experiment Videos
Data mining the protein data bank: automatic detection and assignment of carbohydrate structures.
Thomas Lütteke1, Martin Frank, Claus-W von der Lieth
1Central Spectroscopic Department, German Cancer Research Center, INF 280, D-69120 Heidelberg, Germany. t.luetteke@dkfz.de
Carbohydrate Research
|March 11, 2004
Summary
A new algorithm automatically identifies carbohydrate structures in the Protein Data Bank (PDB), aiding glycomics research. This tool detects over 5,000 glycan chains and identifies errors, improving database quality and data integration.
Area of Science:
- Glycobiology
- Structural Biology
- Bioinformatics
Background:
- Understanding glycoprotein biological roles requires knowledge of glycan 3D structures.
- Locating carbohydrate compounds in the Protein Data Bank (PDB) is challenging due to non-standardized nomenclature.
Purpose of the Study:
- To develop and apply an algorithm for automatic detection and analysis of carbohydrate structures within PDB entries.
- To assess the prevalence of different glycan binding types and identify potential errors in existing PDB data.
Main Methods:
- An algorithm was developed to detect carbohydrate structures using only element types and atom coordinates.
- The algorithm was applied to the PDB to identify and analyze carbohydrate-containing entries and chains.
Main Results:
- The algorithm detected 1663 PDB entries with a total of 5647 carbohydrate chains.
- N-glycosidically bound chains were most frequent, followed by noncovalently bound ligands, with O-glycans being a minority.
- Approximately 30% of carbohydrate-containing PDB entries contained errors.
Conclusions:
- Automatic assignment of carbohydrate structures enhances the integration of glycobiology data with genomic and proteomic resources.
- The algorithm improves PDB database quality by identifying erroneous annotations and structures.
- This approach is crucial for upcoming glycomics projects and advancing our understanding of glycoprotein functions.