Related Experiment Video
Updated: Jun 20, 2025

07:29
Characterization of In Vitro Differentiation of Human Primary Keratinocytes by RNA-Seq Analysis
Published on: May 16, 2020
6.1K
TabDEG: Classifying differentially expressed genes from RNA-seq data based on feature extraction and deep learning
Sifan Feng1, Zhenyou Wang1, Yinghua Jin1
1School of Mathematics and Statistics, Guangdong University of Technology, Guangzhou, Guangdong, China.
Plos One
|July 22, 2024
Summary
This study introduces TabDEG, a novel deep learning model that uses data augmentation to accurately identify differentially expressed genes (DEGs) in small RNA-Seq datasets, improving cancer research.
Area of Science:
- Computational Biology
- Bioinformatics
- Genomics
Background:
- Traditional methods for identifying differentially expressed genes (DEGs) struggle with small sample sizes due to distribution assumptions, leading to high error rates.
- Deep learning (DL) offers a promising alternative for analyzing gene expression data, but challenges remain in labeling and sample size for RNA-Seq data.
- Data augmentation (DA) can generate valuable pseudo-values from limited data, enhancing feature extraction without substantial cost.
Purpose of the Study:
- To develop a robust model, TabDEG, integrating Data Augmentation (DA) with a Deep Learning (DL) framework for improved DEG identification.
- To accurately predict DEGs and their regulatory directions (up-regulation/down-regulation) from gene expression data.
- To address the limitations of traditional models in high-dimensional, small sample size datasets, particularly in cancer genomics.
Main Methods:
- Proposed TabDEG model combining DA and DL-based tabular data modeling.
- Utilized gene expression data from The Cancer Genome Atlas (TCGA) database.
- Compared TabDEG performance against five existing DEG identification methods.
Main Results:
- TabDEG demonstrated high sensitivity and low misclassification rates compared to counterpart methods.
- The model effectively enhances data features for classifying high-dimensional, small sample size datasets.
- Predicted DEGs from TabDEG significantly mapped to important gene ontology terms and cancer-associated pathways.
Conclusions:
- TabDEG is a robust and effective method for identifying DEGs in challenging small sample size datasets.
- The integration of DA and DL provides a powerful approach for analyzing RNA-Seq data in cancer research.
- TabDEG facilitates the discovery of biologically relevant genes and pathways implicated in cancer development.

