Related Experiment Video
Updated: Sep 27, 2025

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
A Data Adaptive Biological Sequence Representation for Supervised Learning.
Hande Cakin1, Berk Gorgulu1, Mustafa Gokce Baydogan1
1Department of Industrial Engineering, Boğaziçi University, İstanbul, Turkey.
This study introduces SW-RF, a novel machine learning approach for analyzing biological sequences. SW-RF effectively represents DNA and protein sequences, improving gene expression prediction and handling missing data in microarrays.
Area of Science:
- Genomics
- Bioinformatics
- Machine Learning
Background:
- Gene expression is crucial for organism function, and DNA microarray technology enables large-scale monitoring.
- Understanding gene regulation mechanisms requires analyzing the relationship between gene expression and nucleotide sequences.
- Identifying local DNA elements (motifs) is key to inferring these regulatory relationships.
Purpose of the Study:
- To propose a novel data-adaptive representation approach for supervised learning on biological sequences.
- To develop a method for predicting biological responses based on sequence data.
- To address challenges in high-dimensionality and missing values common in biological sequence analysis.
Main Methods:
- The study introduces the Sliding Window-Random Forest (SW-RF) method for categorical sequence representation.
- SW-RF represents sequences using overlapping subsequences and a tree-based learner to create a bag-of-words-like representation.
- A lasso logistic regression classifier is trained on the learned representation to identify important patterns.
Main Results:
- The SW-RF approach demonstrated significantly improved accuracy on synthetic and DNA promoter sequence data.
- The method efficiently handles missing values in microarray datasets, a common challenge.
- The learned representation allows for the application of various classifiers for pattern identification.
Conclusions:
- SW-RF provides an effective feature-based representation for categorical sequences, particularly in bioinformatics.
- The approach enhances the accuracy of predicting biological responses from sequence data.
- SW-RF's flexibility makes it applicable to diverse categorical sequence data beyond biological applications.
Related Concept Videos
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Evolutionary Relationships through Genome Comparisons
Genome Annotation and Assembly
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...

