Related Experiment Video
Updated: Apr 19, 2026

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.9K
Evaluating the concordance between sequencing, imputation and microarray genotype calls in the GAW18 data.
Ally Rogers1, Andrew Beck2, Nathan L Tintle1
1Department of Mathematics, Statistics and Computer Science, Dordt College, Sioux Center, IA 51250, USA.
BMC Proceedings
|December 19, 2014
Summary
Genotype errors in Genetic Analysis Workshop 18 (GAW18) data can lead to inaccurate results in genetic studies. Concordance analysis revealed issues with missing data and genotype discrepancies, particularly for low-frequency variants.
Area of Science:
- Genetics
- Bioinformatics
- Statistical Genetics
Background:
- Genotype errors are known to affect the accuracy of genotype-phenotype association studies, potentially increasing false positives or reducing statistical power.
- The Genetic Analysis Workshop 18 (GAW18) dataset lacks a gold standard for genotype calls, necessitating an assessment of data quality.
Purpose of the Study:
- To evaluate the concordance between different genotype calling platforms (sequencing, imputation, microarray) within the GAW18 dataset.
- To identify potential sources and patterns of genotype errors, including missing data and discordance rates.
- To understand the implications of these errors for downstream genetic analyses.
Main Methods:
- Comparative analysis of genotype calls from sequencing, imputation, and microarray data.
- Calculation of missing data rates and genotype concordance rates between platforms.
- Examination of discordance patterns based on minor allele frequency (MAF) and phenotype association.
Main Results:
- High missing data rates were observed for sequenced individuals.
- A modest level of genotype discordance was found between sequencing and imputation platforms, most prevalent in low MAF single-nucleotide polymorphisms (SNPs).
- Some evidence suggested phenotype-specific differences in discordance rates, with instances of conflicting base calls across technologies.
Conclusions:
- Missing genotypes and errors in called genotypes pose a risk of increased Type I errors and reduced power in the downstream analysis of GAW18 data.
- The observed discordance highlights the importance of careful genotype quality control in genetic association studies.
- Findings underscore the need for robust methods to handle genotype uncertainty in large-scale genetic datasets.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
17.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
17.2K
DNA Microarrays
23.3K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
23.3K
Comparing Copy Number Variations and SNPs
19.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.4K
Evolutionary Relationships through Genome Comparisons
7.3K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.3K
RNA-seq
12.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.7K

