评估和减轻稀疏,杂的基因型的隐私风险,通过局部对准与单 haplotype 数据库
Prashant S Emani1,2, Maya N Geradi1,2, Gamze Gürsoy1,2
1Program in Computational Biology and Bioinformatics, Yale University, New Haven, Connecticut 06520, USA.
Genome research
|December 14, 2023
概括
小组遗传数据 (SNP) 存在重新识别的风险. 我们的工具PLIGHT量化了这种隐私泄露,即使是有噪音的数据,并提供了一种消毒方法来保护个人.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 欧米克数据中的单核酸多态 (SNP) 可以重新识别个人及其亲属.
- 以前的研究表明,大型SNP集具有重新识别潜力,但小,杂的基因型集的信息性仍然不清楚.
研究的目的:
- 量化小,杂的SNP集所带来的重新识别风险.
- 从稀疏的基因型数据开发一个用于评估隐私泄露的计算工具.
- 提供一种对基因组数据的消毒方法,以减轻重新识别风险.
主要方法:
- 通过推断跨基因型HMM轨迹 (PLIGHT) 工具套件的隐私泄露的开发.
- 利用基于种群遗传学的隐藏马尔科夫模型 (HMMs) 来对准稀疏的SNP集以引用单元型数据库.
- 分析各种查询场景,包括已知个体,未知个体,环境DNA样本和模拟马赛克.
主要成果:
- 十个常见的,无噪音的SNP足以在约5000个单元类型的数据库中进行个体识别.
- ~20个SNP可以识别两个个体的马赛克中的组件,20-30个可以识别一级亲属.
- PLIGHT从环境样本中识别了使用约30个噪音SNP的个体,并使粗粒度的表型信息泄露成为可能,即使是未列出的个体.
结论:
- 稀疏和杂的SNP集具有显著的重新识别风险.
- 根据有限的基因型数据,PLIGHT有效量化了隐私泄露.
- 开发的净化工具有助于选择性地删除识别SNP,以加强数据隐私,而不需要人口假设.
相关概念视频
Genome-wide Association Studies-GWAS
13.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.5K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K


