Related Experiment Video
Updated: Jun 26, 2026

Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
seqLens: Optimizing Language Models for Genomic Predictions
Mahdi Baghbanzadeh1, Brendan Mann1, Keith A Crandall1
1Computational Biology Institute, Department of Biostatistics and Bioinformatics, Milken Institute School of Public Health, The George Washington University, Washington, DC 20052, USA.
None:
Understanding evolutionary variation in genomic sequences through the lens of language modeling has the potential to revolutionize biological research. Yet to maximize the utility of language modeling in genomics, we must overcome computational challenges in tokenization and model architecture adapted to diverse genomic features across evolutionary timescales. In this study, we investigated key elements in genomic language modeling (gLM), including tokenization, pretraining datasets, fine-tuning approaches, pooling methods, and domain adaptation, and applied the language models to diverse genomic data. We gathered two evolutionarily distinct pretraining datasets: one consisting of 19,551 reference genomes, including over 18,000 prokaryotic genomes (115 B nucleotides) and the remainder eukaryotic genomes, and another more balanced dataset with 1,354 genomes, including 1,166 prokaryotic and 188 eukaryotic reference genomes (180 B nucleotides). We trained five byte-pair encoding tokenizers and pretrained 52 gLMs, systematically comparing different architectures, hyperparameters, and classification heads. We introduce seqLens, a family of models based on disentangled attention with relative positional encoding, which outperforms relatively similar-sized models in 13 of 19 benchmarking phenotypic predictions. We further explore continual pretraining, domain adaptation, and parameter-efficient fine-tuning methods to assess trade-offs between computational efficiency and accuracy. Our findings demonstrate that relevant pretraining data significantly boost performance, alternative pooling techniques can enhance classification, tokenizers with larger vocabulary sizes negatively impact generalization, and gLMs are capable of understanding evolutionary relationships. These insights provide a foundation for optimizing genomic language models for identifying diverse evolutionary genomic features and improving genome annotations.
More Related Videos
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
03:37Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Genomics