Related Experiment Videos
The Pfam protein families database
A Bateman1, E Birney, R Durbin
1The Sanger Centre, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SA, UK. agb@sanger.ac.uk
Nucleic Acids Research
|December 11, 1999
Summary
Pfam is a comprehensive database of protein families, featuring alignments and hidden Markov models. This resource aids in identifying protein domains and understanding proteomes, with version 4.3 covering 1815 families.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein families are fundamental units in understanding protein function and evolution.
- Databases of protein families are essential tools for genomic annotation and analysis.
- Pfam provides curated protein multiple sequence alignments and profile hidden Markov models.
Purpose of the Study:
- To introduce the Pfam database and its latest version (4.3).
- To highlight the utility of Pfam in annotating protein sequences and complete genomes.
- To describe methods for searching genomic DNA against the Pfam library.
Main Methods:
- Collection and curation of protein multiple sequence alignments.
- Development of profile hidden Markov models for each protein family.
- Integration with search tools like Wise2 for genomic DNA analysis.
Main Results:
- Pfam version 4.3 contains 1815 protein families.
- Pfam matches 63% of proteins in SWISS-PROT 37 and TrEMBL 9.
- Pfam identifies up to half of the proteins in complete genomes.
Conclusions:
- Pfam is a valuable resource for protein family identification and annotation.
- The database significantly aids in the analysis of proteomes and genomic data.
- Pfam facilitates the exploration of protein domain architectures across diverse organisms.