Related Experiment Video
Updated: Aug 3, 2026

12:00
A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
Analysis of information content for biological sequences
1EURANDOM, Den Dolech 2, 5612 AZ, Eindhoven, The Chinese Academy of Sciences, Beijing. jzhang@euridice.tue.nl
Summary
This study introduces a new method for analyzing biological sequences by decomposing them into functional domains. The approach improves the identification of critical amino acid residues in proteins.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Biology
Background:
- Identifying functional units in biological sequences requires domain decomposition.
- Current methods involve multi-step processes including sequence alignment and semi-automatic parsing.
- Existing techniques often rely on manual analysis and expert knowledge.
Purpose of the Study:
- To present a novel exploratory approach for parsing and analyzing multiple sequence alignments.
- To improve the identification of functionally important domains and residues within biological sequences.
Main Methods:
- Developed a new method based on analysis-of-variance (ANOVA) decomposition of sequence information content.
- The approach considers composition biases and overdispersion effects among blocks.
- Tested the method on ribosomal protein families.
Main Results:
- The novel approach demonstrated promising performance in analyzing ribosomal protein families.
- It offers an improved way to identify important residues within proteins.
- The method successfully identifies subsets of residues critical to protein function.
Conclusions:
- The ANOVA-based decomposition provides a more effective way to parse biological sequence alignments.
- This method enhances the identification of critical residues, aiding in understanding protein function.
- The approach offers a valuable tool for molecular biology and bioinformatics research.
Related Concept Videos
Phylogenetic Trees
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.The length of the branches can depict time or the relative amount of change among organisms. For instance, the branch length might indicate the number of amino acid changes in the sequence that underlies the...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...

