Related Experiment Video
Updated: Jul 8, 2025

08:03
Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
2.2K
Assessing and mitigating privacy risks of sparse, noisy genotypes by local alignment to haplotype databases
Prashant S Emani1,2, Maya N Geradi1,2, Gamze Gürsoy1,2
1Program in Computational Biology and Bioinformatics, Yale University, New Haven, Connecticut 06520, USA.
Genome Research
|December 14, 2023
Summary
Small sets of genetic data (SNPs) pose reidentification risks. Our tool, PLIGHT, quantifies this privacy leakage, even with noisy data, and offers a sanitization method to protect individuals.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Single nucleotide polymorphisms (SNPs) in omics data can reidentify individuals and their relatives.
- Previous studies demonstrated the reidentification potential of large SNP sets, but the informativeness of small, noisy genotype sets remained unclear.
Purpose of the Study:
- To quantify the reidentification risk posed by small, noisy SNP sets.
- To develop a computational tool for assessing privacy leakage from sparse genotype data.
- To provide a method for sanitizing genomic data to mitigate reidentification risks.
Main Methods:
- Development of the Privacy Leakage by Inference across Genotypic HMM Trajectories (PLIGHT) tool suite.
- Utilizing population-genetics-based hidden Markov models (HMMs) to align sparse SNP sets to reference haplotype databases.
- Analysis of various query scenarios, including known individuals, unknown individuals, environmental DNA samples, and simulated mosaics.
Main Results:
- Ten common, noise-free SNPs were sufficient for individual identification in a database of ~5000 haplotypes.
- ~20 SNPs could identify components in two-individual mosaics, and 20-30 could identify first-order relatives.
- PLIGHT identified individuals using ~30 noisy SNPs from environmental samples, and enabled coarse-grained phenotypic information leakage even for unlisted individuals.
Conclusions:
- Sparse and noisy SNP sets present a significant reidentification risk.
- PLIGHT effectively quantifies privacy leakage from limited genotype data.
- The developed sanitization tool aids in selectively removing identifying SNPs to enhance data privacy without population assumptions.
Related Concept Videos
Genome-wide Association Studies-GWAS
13.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.5K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K

