Related Experiment Video
Updated: Jun 26, 2026

Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
seqLens: Optimizing Language Models for Genomic Predictions
Mahdi Baghbanzadeh1, Brendan Mann1, Keith A Crandall1
1Computational Biology Institute, Department of Biostatistics and Bioinformatics, Milken Institute School of Public Health, The George Washington University, Washington, DC 20052, USA.
Genomic language models (gLMs) can now understand evolutionary relationships in DNA. Optimizing tokenization and pretraining data significantly improves gLM performance for genomic feature identification and annotation.
Area of Science:
- Computational Biology
- Genomics
- Machine Learning
- Evolutionary Biology
Background:
- Language modeling offers a novel approach to understanding genomic sequence variation and evolutionary patterns.
- Existing language models face computational challenges in tokenization and architecture for diverse genomic data across evolutionary scales.
Purpose of the Study:
- To investigate key elements of genomic language modeling (gLM), including tokenization, pretraining datasets, and model architecture.
- To apply and evaluate gLMs for identifying evolutionary genomic features and improving genome annotations.
Main Methods:
- Gathered two distinct pretraining datasets comprising prokaryotic and eukaryotic reference genomes.
- Trained five byte-pair encoding tokenizers and pretrained 52 gLMs, comparing various architectures and hyperparameters.
- Introduced seqLens, a novel model architecture based on disentangled attention with relative positional encoding.
Main Results:
- The seqLens model family demonstrated superior performance in 13 of 19 benchmarking phenotypic predictions compared to similar-sized models.
- Relevant pretraining data significantly enhanced gLM performance, while larger tokenizer vocabularies negatively impacted generalization.
- Alternative pooling techniques improved classification accuracy, and gLMs showed capability in discerning evolutionary relationships.
Conclusions:
- Optimizing tokenization strategies and pretraining datasets is crucial for effective genomic language modeling.
- The developed gLMs can successfully identify diverse evolutionary genomic features and aid in genome annotation.
- Findings provide a foundation for advancing language models in evolutionary genomics research.
More Related Videos
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
03:37Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Genomics