Related Experiment Video
Updated: Jun 20, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Accurate discrimination of conserved coding and non-coding regions through multiple indicators of evolutionary
Matteo Rè1, Graziano Pesole, David S Horner
1Dipartimento di Scienze Biomolecolari e Biotecnologie, Università degli Studi di Milano, Via Celoria 26, 20133 Milano, Italia. matteo.re@unimi.it
This study introduces a machine learning method to distinguish conserved coding sequences from non-coding ones. The approach accurately identifies functional genomic regions, aiding in genome annotation and discovery of novel proteins.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Sequence conservation across genomes indicates functional importance.
- Distinguishing conserved coding from non-coding sequences is crucial, especially with numerous non-coding transcripts.
- Current methods relying on protein databases have limitations in identifying novel proteins and can be affected by annotation errors.
Purpose of the Study:
- To develop a machine learning approach for discriminating conserved coding sequences.
- To enable accurate identification of functional genomic elements without relying on protein similarity searches.
Main Methods:
- A machine learning model, specifically a Support Vector Machine (SVM), was employed.
- The method analyzes evolutionary dynamics statistics of aligned sequences.
- The SVM assigns a probability score indicating whether a sequence is coding or non-coding.
Main Results:
- The developed approach demonstrates high sensitivity and accuracy.
- It effectively discriminates between coding and non-coding sequences.
- The method provides a probability score for classification.
Conclusions:
- The machine learning approach is a sensitive and accurate tool for sequence analysis.
- Applications include identifying conserved coding regions in genomes and differentiating coding from non-coding cDNA.
- This method aids in genome annotation and the discovery of novel functional elements.
More Related Videos
08:23De novo Identification of Actively Translated Open Reading Frames with Ribosome Profiling Data
Published on: February 18, 2022
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Evolutionary Relationships through Genome Comparisons
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Cis-regulatory Sequences