Related Experiment Videos
Comparisons of likelihood and machine learning methods of individual classification
B Guinand1, A Topchy, K S Page
1Department of Fisheries and Wildlife, Michigan State University, East Lansing, MI 48824, USA.
Machine learning classifiers, including artificial neural networks, were compared to traditional likelihood methods for population genetics assignment tests. Likelihood methods are recommended for their accessibility and reliable individual assignment to population of origin.
Area of Science:
- Population Genetics
- Machine Learning
- Bioinformatics
Background:
- Machine learning (ML) classification methods are infrequently applied to population genetic data.
- Traditional methods for assigning individuals to populations (assignment tests) often use parametric likelihood estimations.
- Comparing ML techniques with established methods is crucial for advancing population genetic analyses.
Purpose of the Study:
- To evaluate and compare the accuracy of nonparametric machine learning classifiers against parametric likelihood estimators for population assignment.
- To assess classifier performance across simulated datasets with varying population differentiation (FST), loci number, and allelic diversity.
- To examine classifier performance using empirical lake trout (Salvelinus namaycush) genetic data.
Main Methods:
- Compared artificial neural networks, decision trees, and k-nearest neighbor clustering with likelihood estimation methods.
- Utilized simulated population genetic datasets with controlled parameters (FST, loci, alleles).
- Applied methods to empirical lake trout genetic data exhibiting comparable population differentiation.
Main Results:
- Artificial neural networks and likelihood estimators exhibited lower classification error rates than k-nearest neighbor and decision trees on simulated data.
- Artificial neural networks showed only marginal improvement (0-2.8% lower error) over likelihood methods for simulated data.
- ML classifiers relatively outperformed likelihood estimators on empirical data, indicating an ability to leverage population-specific genotypic patterns.
Conclusions:
- While ML classifiers show promise, particularly on empirical data, likelihood-based methods remain more accessible and reliable for population assignment.
- The complexity in developing and evaluating artificial neural networks makes likelihood methods preferable for routine population genetic studies.
- Further research into ML applications may enhance population genetic analyses, but current likelihood approaches offer a robust and practical solution.
Related Concept Videos
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Comparing the Survival Analysis of Two or More Groups