Related Experiment Video
Updated: Jun 18, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
SIMAP--a comprehensive database of pre-calculated protein sequence similarities, domains, annotations and clusters
Thomas Rattei1, Patrick Tischler, Stefan Götz
1Department of Genome Oriented Bioinformatics, Technische Universität München, Wissenschaftszentrum Weihenstephan, Freising, Germany. t.rattei@wzw.tum.de
The Similarity Matrix of Proteins (SIMAP) database offers pre-calculated protein sequence similarity, aiding function prediction and evolutionary analysis. It now includes expanded sequence data and improved access for large-scale computational biology projects.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein sequence analysis is crucial for predicting function and evolutionary history.
- The exponential growth of protein sequence data makes similarity matrix computation computationally intensive.
- Existing tools struggle to manage the scale of modern sequence databases.
Purpose of the Study:
- To present the Similarity Matrix of Proteins (SIMAP) database as a solution for large-scale protein sequence analysis.
- To highlight SIMAP's comprehensive pre-calculated data and expanded coverage.
- To improve data access and query functionalities for biologists and computational researchers.
Main Methods:
- Pre-calculation of the protein sequence similarity matrix.
- Integration of diverse sequence databases like ENSEMBL and metagenomic data.
- Application of Blast2GO for pre-calculated protein function predictions.
- Development of improved data access via web portal, DAS, and Web-Service.
Main Results:
- SIMAP provides a comprehensive, up-to-date protein sequence similarity matrix covering 48 million proteins (as of Sept 2009).
- Expanded sequence space includes ENSEMBL data and processed metagenomes.
- Pre-calculated protein function predictions and sequence clusters are available.
- Enhanced data access facilitates systematic querying and large-scale downstream projects.
Conclusions:
- SIMAP serves as a vital resource for efficient protein sequence analysis and function prediction.
- The database's comprehensive nature and improved accessibility support advanced computational biology research.
- SIMAP empowers biologists to systematically explore vast sequence spaces for diverse research applications.
Related Concept Videos
Protein Families
Protein Families
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Evolutionary Relationships through Genome Comparisons

