Related Experiment Videos
Selection of conserved blocks from multiple alignments for their use in phylogenetic analysis
1European Molecular Biology Laboratory, Heidelberg, Germany. jose.castresana@embl-heidelberg.de
Molecular Biology and Evolution
|March 31, 2000
Summary
This study presents a new computerized method to improve phylogenetic analysis by removing poorly aligned sequences. This method enhances alignment accuracy and aids in automating large-scale phylogenetic studies.
Area of Science:
- Bioinformatics
- Computational Biology
- Evolutionary Biology
Background:
- Multiple sequence alignments are crucial for phylogenetic analysis.
- Poorly conserved or divergent regions can introduce noise and inaccuracies.
- Manual editing of alignments is time-consuming and difficult to automate.
Purpose of the Study:
- To develop a computerized method for eliminating poorly aligned and divergent regions from multiple sequence alignments.
- To minimize the loss of informative sites during alignment refinement.
- To improve the suitability of alignments for phylogenetic analysis.
Main Methods:
- A novel computerized method was developed to select conserved blocks of positions.
- Selection criteria included contiguous conserved positions, minimal gaps, and conserved flanking regions.
- The method was tested on alignments of 10 mitochondrial proteins from diverse eukaryotic genomes.
Main Results:
- The computerized method effectively removed divergent segments, with higher removal rates in more divergent alignments.
- Post-removal, sequences showed more uniform amino acid composition and reduced pairwise distances.
- Phylogenetic trees derived from refined alignments showed altered topologies, particularly for poorly resolved nodes, supporting animal-fungi clades.
Conclusions:
- The computerized method automates alignment refinement, reducing manual editing needs.
- This facilitates large-scale phylogenetic analyses and improves reproducibility.
- The method enhances the reliability of phylogenetic inferences, especially for complex datasets.