潘多拉:一个工具来估计基因型数据的维度减小稳定性
Julia Haag1, Alexander I Jordan2, Alexandros Stamatakis1,3,4
1Computational Molecular Evolution Group, Heidelberg Institute for Theoretical Studies, Heidelberg 69118, Germany.
Bioinformatics advances
|March 31, 2025
概括
这项研究介绍了Pandora,这是评估基因型数据嵌入的稳定性的新框架. 潘多拉量化了缩小维度分析中的不确定性,提高了人口遗传研究的可靠性.
科学领域:
- 人口遗传学 人口遗传学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 基因型数据集是高维的,对人口结构分析构成挑战.
- 减小尺寸的技术对于推断人口成员至关重要,但需要强大的不确定性估计.
- 目前的方法缺乏标准化的方法来评估基因型数据嵌入的稳定性.
研究的目的:
- 为基因型数据开发稳定性估计框架.
- 量化基因数据的维度缩小固有的不确定性.
- 为验证人口结构推断提供一个工具.
主要方法:
- 开发了Pandora,这是一个基于启动的稳定性估计框架.
- 实施整体嵌入稳定性得分和个人支持值.
- 整合了一个k-means集群方法,用于文化群组分配的不确定性.
主要成果:
- 潘多拉有效量化了嵌入稳定性和个人支持.
- 该框架在模拟和实证基因型数据集上都显示出实用性.
- 潘多拉提供了一种可靠的方法来评估人口遗传分析中的不确定性.
结论:
- 潘多拉为基因型数据维度减少中的不确定性评估提供了一个强大的解决方案.
- 该框架提高了人口结构研究的可靠性和可解释性.
- 潘多拉是公开的,促进了人口遗传学的可复制性研究.
相关概念视频
Comparing Copy Number Variations and SNPs
16.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
16.8K
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K
Genome-wide Association Studies-GWAS
12.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.2K
Evolutionary Relationships through Genome Comparisons
5.6K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.6K
Genetic Variation
241
Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles,...
Genes exist in different versions called alleles,...
241
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


