Related Experiment Video
Updated: Aug 13, 2026

Genome-wide Purification of Extrachromosomal Circular DNA from Eukaryotic Cells
Published on: April 4, 2016
Identification of protein coding regions in genomic DNA
1Department of Molecular, Cellular and Developmental Biology, Universityof Colorado, Boulder 80309-0347, USA.
Abstract:
We have developed a computer program, GeneParser, which identifies and determines the fine structure of protein genes in genomic DNA sequences. The program scores all subintervals in a sequence for content statistics indicative of introns and exons, and for sites that identify their boundaries. This information is weighted by a neural network to approximate the log-likelihood that each subinterval exactly represents an intron or exon (first, internal or last). A dynamic programming algorithm is then applied to this data to find the combination of introns and exons that maximizes the likelihood function. Using this method, we can rapidly generate ranked suboptimal solutions, each of which is the optimum solution containing a given intron-exon junction. We have tested the system on a large collection of human genes. On sequences not used in training, we achieved a correlation coefficient for exon nucleotide prediction of 0.89. For a subset of G + C-rich genes, a correlation coefficient of 0.94 was achieved. We have also quantified the robustness of the method to substitution and frame-shift errors and show how the system can be optimized for performance on sequences with known levels of sequencing errors.
Related Concept Videos
Genomic DNA in Prokaryotes
Genomic Diversity in Bacteria
Although bacterial genomes are much...
Genomic DNA in Eukaryotes
Organization of Genes
From DNA to Protein
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Prokaryotic Gene Structure and Organization

