Related Experiment Video
Updated: Oct 13, 2025

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
Employing phylogenetic tree shape statistics to resolve the underlying host population structure
Hassan W Kayondo1,2, Alfred Ssekagiri3,4, Grace Nabakooza4,5,6
1Institute of Basic Sciences, Technology and Innovation (PAUSTI), Pan African University, Nairobi, Kenya. whkayondo@gmail.com.
This study introduces a novel method using phylogenetic tree statistics and machine learning to accurately identify host population structure. The approach effectively distinguishes between structured and non-structured populations, crucial for understanding disease transmission dynamics.
Area of Science:
- Epidemiology
- Computational Biology
- Population Genetics
Background:
- Host population structure significantly influences infectious disease transmission.
- Phylogenetic trees offer insights into epidemic population structures.
- Identifying population structure guides the selection of appropriate phylogenetic methods.
Purpose of the Study:
- To develop and evaluate a classification procedure for distinguishing between structured and non-structured host populations using phylogenetic tree statistics.
- To assess the performance of machine learning classifiers in identifying population structure from simulated phylogenetic trees.
Main Methods:
- Simulated phylogenetic trees from structured and non-structured host populations.
- Computed eight tree statistics (e.g., number of cherries, Sackin index, Colless index, ladder length).
- Classified trees using Decision Tree (DT), K-nearest neighbor (KNN), and Support Vector Machine (SVM) algorithms, incorporating the basic reproductive number and performing sensitivity analyses.
Main Results:
- Machine learning classifiers achieved high accuracy (AUC > 0.9) in distinguishing between structured and non-structured populations.
- Support Vector Machine (SVM) models demonstrated greater robustness to variations in model parameters and tree size compared to DT and KNN.
- The classification procedure accurately identified a structured population in real-world data using an SVM-polynomial classifier.
Conclusions:
- The developed classification procedure effectively differentiates between structured and non-structured populations.
- SVM classifiers offer superior robustness for population structure analysis.
- The method shows high accuracy when applied to real-world epidemiological data, aiding in the understanding of disease spread.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Phylogenetic Trees
Phylogeny
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Applications of Molecular Taxonomy
Modern Molecular Taxonomy

