KANN:通过最近邻居回归估计遗传祖先的个人资料
Juha Riikonen1, Sini Kerminen1, Aki Havulinna1,2
1Institute for Molecular Medicine Finland, Helsinki Institute of Life Science, University of Helsinki, 00014Helsinki, Finland.
Nucleic acids research
|March 12, 2026
概括
KANN是一种新的,高效的方法来估计大型生物库中的遗传血统. 它使用主要组件并处理连续的祖先资料,优于大型基因组研究的现有工具.
科学领域:
- 遗传学 遗传学 是一个
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 目前的遗传祖先推断方法是计算密集的,限制了它们在大型生物库规模分析中的使用.
- 现有的工具通常需要具有离散人口分配的参考样本,从而限制了它们的适用性.
研究的目的:
- 引入KANN,一种高效的k-最近邻居回归方法,用于个人层面的遗传祖先估计.
- 为了使祖先估计使用主要组件和参考样本与连续的祖先资料.
主要方法:
- 在基因结构的主要组成部分上,KANN利用了k-近邻回归.
- 该方法容纳了在多个来源种群中具有连续祖先概况的参考样本.
主要成果:
- 肯恩的祖先估计与芬兰人口数据的SOURCEFIND方法强烈一致.
- 肯恩在1000个基因组项目的全球多样化的数据集上产生了与ADMIXTURE相似的结果.
结论:
- 肯恩是个体级别遗传祖先估计的高效和有前途的工具.
- 该方法适合进行大规模的基因组研究,因为它具有计算效率和灵活性.
相关概念视频
Evolutionary Relationships through Genome Comparisons
7.2K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.2K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.3K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.3K
Pedigree Analysis
90.4K
Overview
90.4K
Pedigree Analysis
18.8K
18.8K
Genome-wide Association Studies-GWAS
16.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.4K
Residuals and Least-Squares Property
9.7K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.7K

