GRAPE: genomic relatedness detection pipeline.
Alexander Medvedev1,2, Mikhail Lebedev2, Andrew Ponomarev2
1Skolkovo Institute of Science and Technology, Moscow, Russian Federation.
F1000Research
|May 24, 2023
Summary
We developed GRAPE, a Genomic RelAtedness detection PipelinE, to accurately classify genetic relatedness. This open-source tool addresses the need for a fast, reliable solution for genomic data analysis in research and commercial applications.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate relatedness classification is crucial for genomic studies like Genome-Wide Association Studies (GWAS) and genetic linkage analysis.
- Unrecognized population structure can lead to false positives in GWAS, impacting research validity.
- The direct-to-consumer genetic testing market relies on accurate DNA relative matching services.
Purpose of the Study:
- To develop an open-source, end-to-end pipeline for detecting genomic relatedness.
- To provide a fast, reliable, and accurate solution for both close and distant kinship.
- To integrate necessary processing steps for real-world genotypic data and production readiness.
Main Methods:
- Developed GRAPE (Genomic RelAtedness detection PipelinE), an integrated pipeline.
- Incorporated data preprocessing, Identity-by-Descent (IBD) segment detection, and relationship estimation.
- Utilized software development best practices and Global Alliance for Genomics and Health (GA4GH) standards.
Main Results:
- Demonstrated pipeline efficiency on both simulated and real-world genomic datasets.
- GRAPE provides a comprehensive solution for genomic relatedness detection.
- The pipeline is ready for production integration.
Conclusions:
- GRAPE addresses the current gap for an open-source, production-ready relatedness detection pipeline.
- The tool enhances the accuracy and reliability of genomic data analysis.
- GRAPE facilitates advancements in genetic research and commercial applications.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
5.9K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.9K
Genomics
36.6K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.6K
Genome Annotation and Assembly
19.0K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.0K
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K
Genome-wide Association Studies-GWAS
13.7K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.7K


