Distinguishing word identity and sequence context in DNA language models

Melissa Sanabria1, Jonas Hirsch1, Anna R Poetsch2,3

  • 1Biomedical Genomics, Biotechnology Center, Center for Molecular and Cellular Bioengineering, Technische Universitat Dresden, Dresden, Germany.

BMC Bioinformatics
|September 13, 2024
PubMed
Summary

Large language models (LLMs) trained on DNA struggle to learn long-range sequence context when using overlapping k-mers. Further research is needed to understand knowledge representation in biological LLMs.

Related Concept Videos

DNA as a Genetic Template02:05

DNA as a Genetic Template

Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
21.8K
DNA Base Pairing02:27

DNA Base Pairing

Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,
27.2K
Maxam-Gilbert Sequencing01:05

Maxam-Gilbert Sequencing

In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
11.1K
The DNA Helix01:16

The DNA Helix

Overview
138.7K
Labeling DNA Probes03:31

Labeling DNA Probes

DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
8.1K
Nucleic Acid Structure01:25

Nucleic Acid Structure

The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms  a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
6.1K