Related Experiment Videos
Measuring the fit of sequence data to phylogenetic model: allowing for missing data
1Department of Statistics, Department of Biological Sciences, University of South Carolina, Columbia, USA. waddell@stat.sc.edu
Molecular Biology and Evolution
|October 8, 2004
Summary
This study introduces a new method for phylogenetic analysis, enabling accurate model fit assessment even with missing data in molecular sequences. This approach preserves valuable data, improving evolutionary study reliability.
Area of Science:
- Evolutionary biology
- Bioinformatics
- Computational phylogenetics
Background:
- Assessing data-model fit is crucial in phylogenetic and evolutionary studies.
- Standard statistical tests (e.g., G, X²) require complete sequence data, excluding sites with missing information (gaps, ambiguous residues).
- Removing incomplete data sites is often impractical and leads to significant data loss.
Purpose of the Study:
- To develop a method for estimating site-pattern probabilities directly from alignments containing missing data.
- To enable accurate model fit assessment in phylogenetics without discarding incomplete data.
- To extend the applicability of statistical tests to real-world molecular sequence alignments.
Main Methods:
- Utilizing iterative Maximum Likelihood (ML) estimators to compute site-pattern probabilities for columns with missing data.
- Employing Expectation-Maximization (EM) or Newton algorithms for optimization under standard independent and identically distributed (i.i.d.) assumptions.
- Comparing estimated probabilities with model expectations using likelihood-ratio (G) statistics or modified chi-squared (X²) tests.
Main Results:
- Demonstrated that ML estimators can directly handle missing data in sequence alignments.
- Developed a statistically sound approach to calculate G and X² statistics on alignments with missing data.
- Showcased the method's applicability to codon and paired-site models, including those with site-rate variability and Hadamard conjugations.
Conclusions:
- The proposed ML-based method effectively addresses the challenge of missing data in phylogenetic sequence alignments.
- This technique allows for more robust and comprehensive model fit assessment, preserving data integrity.
- The method enhances the utility of phylogenetic and evolutionary analyses by accommodating real-world data complexities.