Related Experiment Video
Updated: Jun 30, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
Detecting conserved coding genomic regions through signal processing of nucleotide substitution patterns.
1Department of Biomolecular Science and Biotechnology, University of Milan, via Celoria 26, 20133 Milan, Italy. matteo.re@unimi.it
Artificial Intelligence in Medicine
|September 24, 2008
Summary
Signal processing techniques can now identify conserved protein-coding regions in genomic sequences. This method analyzes substitution and gap patterns in alignments, improving gene annotation accuracy.
Area of Science:
- Genomics
- Bioinformatics
- Signal Processing
Background:
- Accurate annotation of complete genome sequences, especially protein-coding genes, remains challenging.
- Traditional gene prediction methods struggle with unusual gene structures, and similarity-based approaches may miss novel genes or introduce errors.
- Signal processing techniques, effective at the single genome level, have not yet been applied to comparative genomics for gene identification.
Purpose of the Study:
- To introduce and evaluate the application of signal processing techniques in comparative genomics.
- To develop a novel method for assessing the coding potential of genomic sequence alignments.
- To identify evolutionarily conserved protein-coding regions using signal processing.
Main Methods:
- Developing a signal processing-based test to evaluate coding potential in pairwise genomic sequence alignments.
- Analyzing the pattern and periodicity of substitutions and gaps within alignments.
- Assessing the method's feasibility using annotated human-mouse genomic alignments.
Main Results:
- The proposed signal processing test effectively evaluates the coding potential of genomic alignments.
- The method demonstrates feasibility on human-mouse genomic alignment data.
- Identified conserved protein-coding regions with improved accuracy.
Conclusions:
- Signal processing offers a valuable new approach for comparative genomics.
- The technique aids in identifying evolutionarily conserved protein-coding regions.
- This method enhances the accuracy and scope of gene annotation in comparative genomics.
Related Concept Videos
Multi-species Conserved Sequences
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Single Nucleotide Polymorphisms-SNPs
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...

