Related Experiment Video
Updated: May 31, 2026

Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
A two-stage evolutionary approach for effective classification of hypersensitive DNA sequences
Uday Kamath1, Amarda Shehu, Kenneth A De Jong
1Department of Computer Science, George Mason University, Fairfax, Virginia 20123, USA. ukamath@gmu.edu
This study introduces a two-stage evolutionary algorithm method to improve the accuracy of classifying hypersensitive (HS) sites in DNA sequences. The approach automates Support Vector Machine (SVM) design for better regulatory region identification.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Machine Learning in Biology
Background:
- Hypersensitive (HS) sites are key DNA regulatory regions controlling gene expression.
- Accurate annotation of regulatory regions is crucial for understanding cellular differences and disease pathologies.
- Computational methods, including Support Vector Machines (SVMs), are used to identify HS sequences.
Purpose of the Study:
- To propose an automated method for designing Support Vector Machines (SVMs) to enhance the classification accuracy of DNA sequences.
- To improve the identification of regulatory regions by optimizing SVM design for hypersensitive (HS) site detection.
Main Methods:
- A two-stage method employing evolutionary algorithms for SVM design automation.
- Stage 1: An evolutionary algorithm designs optimal sequence motifs for feature vector creation.
- Stage 2: A second evolutionary algorithm optimizes SVM kernel functions and parameters for HS/non-HS classification.
Main Results:
- The proposed two-stage evolutionary algorithm method significantly improves SVM classification accuracy for identifying HS sites.
- The automated SVM design process leads to more precise discrimination between regulatory and non-regulatory DNA sequences.
Conclusions:
- This method offers a significant advancement in automating the analysis of biological sequences, particularly for regulatory region identification.
- The approach is broadly applicable to biological sequence analysis tasks, with source code made publicly available.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy
Methods of Classification and Identification
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Sanger Sequencing
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
