Related Experiment Video
Updated: Feb 4, 2026

Pattern-based Search of Epigenomic Data Using GeNemo
Published on: October 8, 2017
Identity inference of genomic data using long-range familial searches
Yaniv Erlich1,2,3,4, Tal Shor5, Itsik Pe'er2,3
1MyHeritage, Or Yehuda 6037606, Israel. erlichya@gmail.com.
Consumer genomics databases enable law enforcement to identify suspects through distant relatives. This study shows 60% of European descent individuals could be identified, raising privacy concerns.
Area of Science:
- Genomic privacy
- Forensic genealogy
- Bioinformatics
Background:
- Consumer genomics databases now contain millions of individuals' genetic information.
- Law enforcement agencies are increasingly using these databases for suspect identification through familial matching.
Purpose of the Study:
- To investigate the effectiveness and reach of using consumer genomics databases for identifying individuals.
- To assess the potential for implicating a significant portion of the population, particularly those of European descent.
Main Methods:
- Analysis of genomic data from 1.28 million individuals tested via consumer genomics.
- Projection of identification success rates based on familial matching thresholds (e.g., third-cousin or closer).
Main Results:
- Approximately 60% of searches for individuals of European descent are projected to yield a third-cousin or closer match.
- The technique has the potential to identify nearly any U.S. individual of European descent in the near future.
- Research participants in public sequencing projects can also be identified using this method.
Conclusions:
- The widespread use of consumer genomics poses significant privacy risks, enabling identification through familial DNA.
- A mitigation strategy and policy implications for human subject research are proposed in response to these findings.
More Related Videos
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
12:39A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
Related Concept Videos
Protein Families
Protein Families
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Gene Families
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Genomics