GtTR: Bayesian estimation of absolute tandem repeat copy number using sequence capture and high throughput
Devika Ganesamoorthy1, Minh Duc Cao1, Tania Duarte1
1Institute for Molecular Biosciences, University of Queensland, Brisbane, Australia.
BMC Bioinformatics
|July 18, 2018
Summary
We developed GtTR, a novel Bayesian algorithm for genotyping tandem repeats, enabling population-scale analysis of genomic variation. This cost-effective method improves variant discovery in complex diseases.
Area of Science:
- Genomics
- Bioinformatics
- Human Genetics
Background:
- Tandem repeats are crucial genomic regions prone to variation, contributing to human diversity and complex diseases.
- Analyzing tandem repeats is challenging due to technical limitations in high-throughput sequencing.
- Genomic variation in tandem repeats impacts coding and regulatory regions.
Purpose of the Study:
- To develop a novel, cost-effective targeted sequencing approach for simultaneous analysis of hundreds of tandem repeats.
- To create a Bayesian algorithm (GtTR) for population-scale genotyping of tandem repeats.
- To improve the understanding of tandem repeat variation in complex diseases.
Main Methods:
- Developed a Bayesian algorithm (GtTR) combining long-read reference data with short-read counting.
- Utilized PacBio long-read sequencing for reference dataset generation.
- Validated genotyping accuracy using PCR sizing analysis.
Main Results:
- GtTR achieved high accuracy in estimating VNTR copy number, with improved performance using PCR reference data.
- Genotype resolution increased with sequencing depth, showing significant improvements at higher coverage.
- Sequencing-based genotype estimates showed strong correlation with PCR validation results.
Conclusions:
- The novel GtTR approach offers a cost-effective method for exploring tandem repeat variation in Genome-Wide Association Studies (GWAS).
- This method facilitates the discovery of previously unrecognized repeat variations relevant to complex diseases.
- Enhancing the accuracy of the reference dataset can further improve genotyping precision.
Related Concept Videos
Cis-regulatory Sequences
11.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.9K
Cis-regulatory Sequences
4.2K
4.2K
Sequences
280
Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
280
Sanger Sequencing
774.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
774.8K
Arithmetic Sequences
240
An arithmetic sequence is a structured arrangement of numbers where each term is derived by adding a constant value, known as the common difference, to the previous term. This consistent pattern allows for the efficient computation of any term within the sequence as well as the cumulative sum of multiple terms. The formula for finding the nth term of an arithmetic sequence is:Here, aₙ represents the nth term of the sequence, a is the first term, d is the common difference, and n is the...
240
Next-generation Sequencing
98.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
98.7K


