Related Experiment Videos
Detecting the number of clusters of individuals using the software STRUCTURE: a simulation study
G Evanno1, S Regnaut, J Goudet
1Department of Ecology and Evolution, Biology building, University of Lausanne, CH 1015 Lausanne, Switzerland.
Molecular Ecology
|June 23, 2005
Summary
Identifying genetically homogeneous groups is key in population genetics. A new method using STRUCTURE software, with a DeltaK statistic, accurately identifies the highest level of genetic structure, even with uneven population dispersal.
Area of Science:
- Population genetics
- Computational biology
- Bioinformatics
Background:
- Identifying genetically homogeneous groups is a fundamental challenge in population genetics.
- The Bayesian algorithm in STRUCTURE software is widely used for this purpose.
- Its effectiveness with non-homogeneous dispersal patterns remains untested.
Purpose of the Study:
- To evaluate the STRUCTURE algorithm's ability to detect the correct number of genetic clusters (K).
- To test this under various non-homogeneous population dispersal scenarios.
- To assess the utility of the 'log probability of data' and a novel DeltaK statistic.
Main Methods:
- Utilized an individual-based model to generate genetic data under diverse dispersal scenarios.
- Applied the STRUCTURE Bayesian algorithm to analyze simulated datasets.
- Calculated the 'log probability of data' and the ad hoc DeltaK statistic for successive K values.
Main Results:
- The standard 'log probability of data' often failed to accurately estimate the true number of clusters (K).
- The DeltaK statistic effectively identified the uppermost hierarchical level of genetic structure across tested scenarios.
- Results demonstrated sensitivity to marker type (AFLP vs. microsatellite), loci number, populations sampled, and individuals typed.
Conclusions:
- The DeltaK statistic offers a more reliable method for inferring the number of genetic clusters with STRUCTURE.
- Accurate detection of population structure is feasible even with complex dispersal patterns.
- Optimal results depend on careful consideration of genetic marker choice and data sample size.