Related Experiment Video
Updated: Apr 16, 2026

A Strategy to Identify de Novo Mutations in Common Disorders such as Autism and Schizophrenia
Published on: June 15, 2011
The distribution and mutagenesis of short coding INDELs from 1,128 whole exomes
Danny Challis1,2,3, Lilian Antunes4,5,6, Erik Garrison7
1Human Genome Sequencing Center, Baylor College of Medicine, Houston, TX, 77030, USA. dannychallis@gmail.com.
Background:
Identifying insertion/deletion polymorphisms (INDELs) with high confidence has been intrinsically challenging in short-read sequencing data. Here we report our approach for improving INDEL calling accuracy by using a machine learning algorithm to combine call sets generated with three independent methods, and by leveraging the strengths of each individual pipeline. Utilizing this approach, we generated a consensus exome INDEL call set from a large dataset generated by the 1000 Genomes Project (1000G), maximizing both the sensitivity and the specificity of the calls.
Results:
This consensus exome INDEL call set features 7,210 INDELs, from 1,128 individuals across 13 populations included in the 1000 Genomes Phase 1 dataset, with a false discovery rate (FDR) of about 7.0%.
Conclusions:
In our study we further characterize the patterns and distributions of these exonic INDELs with respect to density, allele length, and site frequency spectrum, as well as the potential mutagenic mechanisms of coding INDELs in humans.
Related Concept Videos
Exon Recombination
Exon shuffling follows “splice frame rules.” Each exon...
Mutations
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Mutations
Spontaneous and Induced Mutations
In vitro Mutagenesis
Genome Copying Errors

