Related Experiment Video
Updated: Jun 13, 2025

16:24
Analyzing and Building Nucleic Acid Structures with 3DNA
Published on: April 26, 2013
20.5K
Distinguishing word identity and sequence context in DNA language models
Melissa Sanabria1, Jonas Hirsch1, Anna R Poetsch2,3
1Biomedical Genomics, Biotechnology Center, Center for Molecular and Cellular Bioengineering, Technische Universitat Dresden, Dresden, Germany.
BMC Bioinformatics
|September 13, 2024
Summary
Large language models (LLMs) trained on DNA struggle to learn long-range sequence context when using overlapping k-mers. Further research is needed to understand knowledge representation in biological LLMs.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Transformer-based large language models (LLMs) show promise for analyzing biological sequences due to their natural language processing capabilities.
- Tokenization allows LLMs to learn complex relationships within biological sequences, similar to words in natural language.
- Masked token prediction during training enables LLMs to capture both local token identity and broader sequence context.
Purpose of the Study:
- To develop and apply methodologies for interrogating the learning processes of biological LLMs.
- To evaluate the interpretability and task-specific potential of LLMs trained on genomic data.
- To assess the capability of DNA language models to learn sequence context independent of tokenization strategies.
Main Methods:
- Utilized DNABERT, a DNA language model trained on the human genome using overlapping k-mers as tokens.
- Interrogated model predictions and extracted token embeddings to understand learned representations.
- Developed a novel fine-tuning task to predict subsequent tokens of varying lengths without overlap, isolating context learning.
Main Results:
- The DNABERT model with overlapping k-mers demonstrated limitations in learning extended sequence context.
- Learned embeddings primarily reflected token sequence identity rather than broader contextual information.
- Despite context learning challenges, the model achieved strong performance on genome-biology-specific fine-tuning tasks.
Conclusions:
- Overlapping k-mer tokenization in biological LLMs may hinder the learning of long-range sequence dependencies.
- LLMs with overlapping tokens are suitable for tasks where local token features are paramount and long-range context is less critical.
- There is a critical need for robust methods to interrogate and understand knowledge representation within biological LLMs.
Related Concept Videos
DNA as a Genetic Template
21.8K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
21.8K
DNA Base Pairing
27.2K
Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,
27.2K
Maxam-Gilbert Sequencing
11.1K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.1K
The DNA Helix
138.7K
Overview
138.7K
Labeling DNA Probes
8.1K
DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
8.1K
Nucleic Acid Structure
6.1K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
6.1K

