Related Experiment Video
Updated: Aug 24, 2025

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
A roadmap to robust discriminant analysis of principal components
Catherine Cullingham1, Rhiannon M Peery1, Joshua M Miller2
1Department of Biology, Carleton University, Ottawa, Ontario, Canada.
Abstract:
Identification of population structure is a common goal for a variety of applications, including conservation, wildlife management, and medical genetics. The outcome of these analyses can have far reaching implications; therefore, it is important to ensure an understanding of the strengths and weaknesses of the methodologies used. Increasing in popularity, the discriminant analysis of principal components (DAPC) method incorporates combinations of genetic variables (alleles) into a model that differentiates individuals into genetic clusters. However, users may not have a full understanding of how to best parameterize the model. In this issue of Thia (Molecular Ecology Resources, 2022) looks under the hood of the DAPC. Using simulated data, he demonstrates the importance of careful parameter selection in developing a DAPC model, what the implications are for over-fitting the model, and finally, how best to evaluate the results of DAPC models. This work highlights the issues that can arise when over-parameterizing the DAPC model when gene flow is high among clusters and provides important guidelines to ensure researchers are making conclusions that are biologically relevant.
Related Concept Videos
Vector Algebra: Method of Components
In many applications, the magnitudes and directions of...
Three-Dimensional Analysis of Strain
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Principal Moments of Area
The principal moment of inertia axes are the...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Factorial Design

