Related Experiment Video
Updated: May 3, 2026

Hi-C: A Method to Study the Three-dimensional Architecture of Genomes.
Published on: May 7, 2010
Low hanging fruit: a subset of human cSNPs is both highly non-uniform and predictable
Monica M Horvath1, John W Fondon, Harold R Garner
1McDermott Center for Human Growth and Development, The University of Texas Southwestern Medical Center, 5323 Harry Hines Blvd, Dallas, TX 75390-8591, USA. monica.horvath@utsouthwestern.edu
Abstract:
We present a point mutation classification method that contrasts SNP databases and has the potential to illuminate the relative mutational load of genes caused by codon bias. We group point variation gleaned from public databases by their wild-type and mutant codons, e.g. codon mutation classes (CMCs, 576 possible such as ACG-->ATG), whose frequencies in a database are assembled into a BLOSUM-style matrix describing the likelihood of observing all possible single base codon changes as tuned by the intertwined effects of mutation rate and selection. The rankings of the CMCs in any database are reshuffled according to the population stratification of the typical genotyping experiment producing that resource's data. Analysis of four independent databases reveals that a considerable fraction of mutation in functional genes can be described by a few CMCs regardless of gene identity or population stratification in the genotyping experiment. For example, the top 5% (29/576) of CMCs account for 27.4% of the observed variants in dbSNP while the bottom 5% account for only 0.02%. For non-synonymous disease-causing mutation, 40.8% are described by the top 5% of all possible non-silent CMCs (22/438). Overall, the most observed polymorphism is a G-->A transition at CpG dinucleotides causing ACG, TCG, GCG, and CCG to frequently undergo silent mutation in any gene due to the putative lack of impact on the protein product. In order to assess how well CMC spectrums estimate the aggregate non-synonymous mutational trends of a single gene, a CMC matrix was applied to seven unrelated genes to compute the most likely point mutations. In excess of 87% of these mutation predictions are historically known to play an important role in a disease state according to published literature. CMC-based mutation prediction may aid design and execution of direct association genotyping studies.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Evolutionary Relationships through Genome Comparisons
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Single Nucleotide Polymorphisms-SNPs
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...

