Related Experiment Video
Updated: Jul 8, 2026

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
scGenoByte: a GenoByte embedding transformer with biological priors for cell type annotation
Jiongsen Yao1, Yong Xu1, Jinjin Ma2
1School of Computer Science and Engineering, South China University of Technology, Building B3, 382 Waihuan East Road, Guangzhou Higher Education Mega Centre, Panyu District, Guangzhou 510006, Guangdong, China.
scGenoByte enhances cell representation learning in single-cell RNA sequencing (scRNA-seq) by modeling the full transcriptome using biologically informed GenoBytes. This approach improves cell annotation and analysis of cellular heterogeneity.
Area of Science:
- Computational Biology
- Genomics
- Bioinformatics
Background:
- Effective cell representation learning is vital for single-cell RNA sequencing (scRNA-seq) analysis, including cell annotation and understanding cellular heterogeneity.
- Current foundation models outperform traditional methods but often discard gene information due to data sparsity and model complexity.
- Modeling the complete transcriptome for cell representation remains a significant computational challenge.
Purpose of the Study:
- To introduce scGenoByte, a unified framework for enhanced cell representation learning via biologically informed full-gene modeling.
- To address the limitations of existing methods that compromise gene information by proposing an efficient full-transcriptome modeling approach.
- To integrate biological priors for improved cell function and representation analysis.
Main Methods:
- Developed GenoBytes, biologically coherent units leveraging protein-protein interaction and gene paralogy networks for efficient full transcriptome modeling.
- Integrated protein representations with GenoByte embeddings to capture critical protein information.
- Employed pathway activity prediction as an auxiliary task for pathway-guided regularization.
Main Results:
- scGenoByte demonstrated superior performance across eight diverse scRNA-seq datasets compared to existing methods.
- The framework effectively models the full transcriptome by incorporating biological priors.
- The integration of biological information significantly enhances cell representation learning.
Conclusions:
- scGenoByte offers an effective solution for computationally challenging full transcriptome modeling in scRNA-seq.
- The study confirms the efficacy of combining full-gene context with biological priors for improved cell representation.
- The framework advances cell annotation and the deciphering of cellular heterogeneity.
Related Concept Videos
Genome Annotation and Assembly
Cell Specific Gene Expression
Cell Specific Gene Expression
Genomics
Background and Environment Affect Phenotype
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
Methods of Nuclear Reprogramming
