Related Experiment Video
Updated: Jul 10, 2026

Transcriptomic Analysis of C. elegans RNA Sequencing Data Through the Tuxedo Suite on the Galaxy Project
Published on: April 8, 2017
GenART: Reading the genome's language with adaptive "words"
Kehan Chen1, Ping Han1, Zhe Lin2
1School of Informatics, National Institute for Data Science in Health and Medicine, Xiamen University, Xiamen 361000, China.
None:
Tokenization is both a prerequisite and a central challenge for genomic language models, due to the inherent difficulty of delineating meaningful segments within continuous DNA sequences. We present GenART, a genomic language model framework that dynamically segments DNA into variable-length words directly from raw sequences without manual annotation. Its key adaptive tokenization module infers biologically meaningful boundaries from multi-scale contextual signals during pretraining. GenART demonstrates competitive performance across diverse tasks, outperforming leading approaches. Word boundaries derived from unsupervised tokenization show strong correspondence with base-resolution regulatory signals, such as DNA methylation, and with multi-nucleotide functional elements annotated in GENCODE v.49. These findings highlight adaptive tokenization as a powerful strategy for building interpretable and biologically responsive genomic language models.
More Related Videos
Related Concept Videos
Genetic Lingo
Genomics
Genome Size and the Evolution of New Genes
Genome Size and the Evolution of New Genes
DNA as a Genetic Template
DNA as a Genetic Template

