Related Experiment Video
Updated: Jun 23, 2026

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
A new effective method for estimating missing values in the sequence data prior to phylogenetic analysis
Abdoulaye Baniré Diallo1, François-Joseph Lapointe, Vladimir Makarenkov
1Département d'informatique, Université du Québec à Montréal, C.P. 8888, Succ. Centre-Ville, Montréal (Québec), H3C 3P8, Canada. banire(at)lacim.uqam.ca
This study introduces Probabilistic Estimation of Missing Values (PEMV) for more accurate phylogenetic inference from nucleic acid data. PEMV outperforms existing methods by estimating missing nucleotides before calculating evolutionary distances.
Area of Science:
- Bioinformatics
- Computational Biology
- Evolutionary Biology
Background:
- Phylogenetic inference from nucleic acid data is crucial for understanding evolutionary relationships.
- Missing bases in sequence data pose a significant challenge to accurate phylogenetic analysis.
- Existing methods for handling missing data can lead to inaccuracies in tree reconstruction.
Purpose of the Study:
- To introduce a novel method, Probabilistic Estimation of Missing Values (PEMV), for handling missing bases in phylogenetic inference.
- To evaluate the performance of PEMV against established methods.
- To improve the accuracy of evolutionary distance calculations and subsequent phylogenetic tree construction.
Main Methods:
- Developed a probabilistic approach for estimating missing nucleotides based on Jukes-Cantor and Kimura 2-parameter models.
- Compared PEMV with "Ignoring Missing Sites" (IMS) and "Proportional Distribution of Missing and Ambiguous Bases" (PDMAB) within PAUP.
- Assessed method performance using simulations with SeqGen and Bio NJ, and analyzed a real dataset of eutherian mammals.
Main Results:
- PEMV demonstrates improved accuracy in phylogenetic inference compared to IMS and PDMAB.
- Simulations confirmed the enhanced performance of PEMV.
- Analysis of the eutherian mammal dataset also showed advantages of the PEMV approach.
Conclusions:
- Probabilistic Estimation of Missing Values (PEMV) offers a more accurate approach to phylogenetic inference with incomplete nucleic acid data.
- The method provides a robust framework for estimating evolutionary relationships when sequence data contains missing bases.
- PEMV represents a significant advancement in handling missing data for molecular phylogenetics.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Phylogenetic Trees
Phylogenetic Trees
Microbial Phylogeny
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Gene Evolution - Fast or Slow?
In contrast, regions which code...

