Related Experiment Videos
Assessing strategies for improved superfamily recognition
Ian Sillitoe1, Mark Dibley, James Bray
1Biomolecular Structure and Modelling Unit, Department of Biochemistry and Molecular Biology, University College London, UK.
Summary
We developed methods to map protein structures onto genome sequences, enabling annotation of over 70% of genes in 120 genomes using Hidden Markov Model (HMM) libraries.
Area of Science:
- Genomics
- Structural Biology
- Bioinformatics
Background:
- Public repositories contain millions of genome sequences and thousands of protein structures.
- Mapping structural data to genomic sequences is crucial for understanding protein function.
- Sequence-based methods are increasingly powerful for inferring structure from sequence.
Purpose of the Study:
- To review and assess sequence-based strategies for structural annotation of genomes.
- To describe the protocol for providing CATH structural annotations for completed genomes.
- To evaluate Hidden Markov Model (HMM) technologies for protein superfamily recognition.
Main Methods:
- Utilized Hidden Markov Model (HMM) technologies for superfamily recognition.
- Developed and tested SAMOSA (sequence augmented models of structure alignments) models.
- Employed a CATH-ISL (CATH-Integrated Structural Library) expanded HMM library.
- Annotated protein sequences from 120 genomes across three kingdoms.
Main Results:
- A single-seed HMM library recognized 76% of remote homologs.
- SAMOSA models showed minimal gain in homolog recognition but improved alignment quality for very remote homologs.
- An expanded 1D-HMM library (CATH-ISL) increased coverage to 86%.
- Up to 70% of genes in 120 annotated genomes were assigned to CATH superfamilies.
- Recruitment of sequences expanded the CATH database eightfold.
Conclusions:
- Sequence-based methods, particularly HMMs and CATH-ISL, are effective for structural annotation of genomes.
- These methods enable large-scale assignment of genes to protein structural superfamilies.
- The approach significantly expands protein domain databases and aids in understanding genomic content.