相关实验视频
Updated: Jun 13, 2025

16:24
Analyzing and Building Nucleic Acid Structures with 3DNA
Published on: April 26, 2013
20.5K
在DNA语言模型中区分单词身份和序列上下文
Melissa Sanabria1, Jonas Hirsch1, Anna R Poetsch2,3
1Biomedical Genomics, Biotechnology Center, Center for Molecular and Cellular Bioengineering, Technische Universitat Dresden, Dresden, Germany.
BMC bioinformatics
|September 13, 2024
概括
在DNA上训练的大型语言模型 (LLM) 在使用重叠的k-mers时,难以学习远程序列上下文. 需要进一步的研究,以了解生物LLMs的知识表示.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 基于变压器的大型语言模型 (LLM) 显示出由于它们的自然语言处理能力,对分析生物序列的承诺.
- 代币化允许LLM学习生物序列中的复杂关系,类似于自然语言中的单词.
- 在培训期间,掩盖的令牌预测使LLM能够捕获本地令牌身份和更广泛的序列上下文.
研究的目的:
- 开发和应用用于询问生物LLMs的学习过程的方法.
- 评估在基因组数据上训练的LLM的可解释性和特定任务潜力.
- 评估DNA语言模型学习序列上下文的能力,独立于代币化策略.
主要方法:
- 利用DNABERT,一种在人类基因组上训练的DNA语言模型,使用重叠的k-mers作为令牌.
- 询问模型预测和提取的符号嵌入,以理解学习的表示.
- 开发了一种新的微调任务,以预测不同长度的后续令牌,而不会重叠,隔离上下文学习.
主要成果:
- 具有重叠k-mers的DNABERT模型在学习扩展序列上下文方面表现出局限性.
- 学习的嵌入主要反映了代币序列身份,而不是更广泛的上下文信息.
- 尽管存在上下文学习挑战,但该模型在基因组生物学特定的微调任务上取得了强的表现.
结论:
- 在生物LLM中重叠的k-mer标记化可能会阻碍长距离序列依赖的学习.
- 具有重叠令牌的LLM适用于当地的令牌特征至关重要,远程环境不那么关键的任务.
- 迫切需要强大的方法来询问和理解生物LLM中的知识表示.
相关概念视频
DNA as a Genetic Template
21.8K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
21.8K
DNA Base Pairing
27.2K
Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,
27.2K
Maxam-Gilbert Sequencing
11.1K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.1K
The DNA Helix
138.7K
Overview
138.7K
Labeling DNA Probes
8.1K
DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
8.1K
Nucleic Acid Structure
6.1K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
6.1K

