Related Experiment Video
Updated: May 17, 2025

05:37
Single-Molecule Fluorescence Visualization of DNA Polymerase Dynamics at G-Quadruplexes
Published on: April 4, 2025
300
Benchmarking DNA large language models on quadruplexes
Oleksandr Cherednichenko1, Alan Herbert1,2, Maria Poptsova1
1International Laboratory of Bioinformatics, HSE University, Moscow, Russia.
Computational and Structural Biotechnology Journal
|March 31, 2025
Summary
This study benchmarks large language models (LLMs) for whole-genome G-quadruplex (GQ) annotation, finding that different LLM architectures excel at detecting distinct functional genomic elements.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Large language models (LLMs) show promise in predicting genomic elements.
- Selecting the optimal LLM for specific tasks like whole-genome annotation remains challenging.
- LLMs in genomics are broadly categorized into transformer-based, long convolution-based, and state-space models (SSMs).
Purpose of the Study:
- To benchmark different large language model (LLM) architectures for whole-genome annotation of G-quadruplexes (GQ).
- To evaluate the performance of transformer-based, long convolution-based, and state-space models in identifying GQ structures.
- To determine which LLM architectures are best suited for specific downstream genomic tasks.
Main Methods:
- Benchmarking three LLM architectures (transformer-based, long convolution-based, SSMs) for whole-genome G-quadruplex (GQ) mapping.
- Evaluating model performance using F1 and Matthews Correlation Coefficient (MCC) metrics.
- Analyzing whole-genome annotations to identify distinct functional elements recovered by each model type.
Main Results:
- All evaluated LLMs performed comparably, with DNABERT-2 and HyenaDNA showing superior F1 and MCC scores.
- HyenaDNA demonstrated enhanced recovery of quadruplexes in distal enhancers and intronic regions.
- Different LLM architectures, particularly HyenaDNA and Caduceus versus transformer-based models, showed distinct patterns in de novo quadruplex generation.
Conclusions:
- LLM architectures with varying context lengths can detect distinct functional regulatory elements.
- The choice of LLM architecture is crucial for specific genomic tasks, as different models offer complementary strengths.
- This study highlights the importance of selecting appropriate LLMs for accurate and comprehensive whole-genome annotation.
Related Concept Videos
Maxam-Gilbert Sequencing
10.6K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
10.6K
Ribosome Profiling
3.4K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.4K

