Related Experiment Videos
Quantitative assessment of relationship between sequence similarity and function similarity
1Digital Biology Laboratory, Department of Computer Science and Christopher S, Bond Life Sciences Center, University of Missouri, Columbia, Missouri 65211, USA. joshitr@missouri.edu <joshitr@missouri.edu>
BMC Genomics
|July 11, 2007
Summary
Sequence similarity is a key step in protein annotation, but can lead to errors. This study benchmarks sequence-function relationships to improve the accuracy of automated protein function assignment.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Comparative sequence analysis is crucial for genome annotation.
- Sequence similarity can lead to errors in protein function assignment.
- Systematic analysis of sequence-based function assignment quality is important.
Purpose of the Study:
- To analyze the relationship between sequence and function similarity in proteins.
- To quantify the correlation between functional and sequence similarity.
- To compare these correlations against random protein pairs.
Main Methods:
- Analysis of proteins from four model organisms (Arabidopsis thaliana, Saccharomyces cerevisiae, Caenorrhabditis elegans, Drosophila melanogaster).
- Functional similarity measured using Gene Ontology (GO) classifications (biological process, molecular function, cellular component).
- Sequence similarity assessed by sequence identity and alignment statistical significance.
Main Results:
- Identified various sequence-function relationships using different methods (BLAST, PSI-BLAST, sequence identity, Expectation Value).
- Compared GO indices against semantic similarity approaches.
- Examined relationships within and between genomes for all three GO categories.
Conclusions:
- Established a benchmark for estimating confidence in sequence-based function assignment.
- Highlighted the importance of systematic analysis to mitigate errors in annotation.
- Provided insights into the nuances of sequence-function relationships across different comparison types.
Related Concept Videos
Modern Molecular Taxonomy
511
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
511
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K
Protein Families
16.5K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.5K
Convergent Evolution
31.2K
Evolution shapes the features of organisms over time, ensuring that they are suited for the environments in which they live. Sometimes, selection pressure leads to the rise of similar but unrelated adaptations in organisms with no recent common ancestors, a process known as convergent evolution.
31.2K
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Multi-species Conserved Sequences
4.6K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.6K