Related Experiment Video
Updated: Mar 13, 2026

Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
KANN: estimation of genetic ancestry profiles by nearest neighbor regression
Juha Riikonen1, Sini Kerminen1, Aki Havulinna1,2
1Institute for Molecular Medicine Finland, Helsinki Institute of Life Science, University of Helsinki, 00014Helsinki, Finland.
Abstract:
State-of-the-art methods for inferring individual-level genetic ancestry are based on statistical models for haplotype data. Unfortunately, these methods are computationally demanding, making them impractical for biobank-scale analyses. In this paper, we describe KANN, an efficient k-nearest neighbor regression method for individual-level ancestry estimation with respect to predefined source populations using only principal components of genetic structure. Contrary to the existing tools that can only use reference samples with discrete source population assignment, KANN enables the use of reference samples with continuous ancestry profiles across multiple source populations. We observe that KANN's ancestry estimates agree well with the haplotype-based method SOURCEFIND when estimating ancestry profiles across up to 10 Finnish source populations on a dataset of 18 125 Finnish samples from THL Biobank. In the 1000 Genomes Project data containing globally diverse genetic backgrounds, KANN produces highly similar results to the ADMIXTURE software. Based on our results, KANN is a promising tool for ancestry estimation in large-scale genomic studies.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Pedigree Analysis
Pedigree Analysis
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
