Related Experiment Video
Updated: Jun 22, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Statistical assessment of discriminative features for protein-coding and non coding cross-species conserved sequence
Teresa M Creanza1, David S Horner, Annarita D'Addabbo
1Istituto di Studi sui Sistemi Intelligenti per l'Automazione, CNR, Via Amendola 122/D-I, Bari, Italy. creanza@ba.issia.cnr.it
Identifying protein-coding elements in conserved sequences is challenging. Comparative features, particularly those measuring nucleotide substitutions, are most effective for distinguishing coding from non-coding DNA, improving prediction accuracy.
Area of Science:
- Molecular Biology
- Bioinformatics
- Genomics
Background:
- Distinguishing protein-coding elements within conserved mammalian sequences presents a significant research challenge.
- Numerous features have been proposed to automate this classification, necessitating a thorough statistical evaluation of their efficacy.
Purpose of the Study:
- To systematically assess and compare the statistical differences between coding and non-coding conserved sequences.
- To evaluate the predictive accuracy of various features and classifiers for discriminating these sequence types.
Main Methods:
- Comparative analysis of feature distributions between coding and non-coding conserved sequences.
- Development and testing of prediction models using single and combined features.
- Analysis of feature performance across varying sequence lengths and species (human, mouse, rat).
Main Results:
- All analyzed features showed statistically significant differences between coding and non-coding sequences.
- The proportion of synonymous nucleotide substitutions per synonymous site emerged as the most powerful discriminant feature.
- Classifiers utilizing comparative features generally outperformed those using intrinsic features.
- Combining comparative and intrinsic features significantly enhanced prediction accuracy, irrespective of sequence length.
Conclusions:
- Distinct patterns were observed for individual and combined feature sets in classifying coding and non-coding elements.
- Comparative features demonstrated higher accuracy in classifying coding sequences, likely due to their sensitivity to evolutionary constraints imposed by the genetic code.
More Related Videos
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...