相关实验视频
Updated: Jul 2, 2025

03:45
A Rapid Method to Confine and Safely Handle Bees in the Field
Published on: August 23, 2024
758
DNABERT-S: 开拓物种分化与物种意识的DNA嵌入
Zhihan Zhou1, Weimin Wu1, Harrison Ho2,3
1Department of Computer Science, Northwestern University.
ArXiv
|February 27, 2024
概括
DNABERT-S是一种新的基因组模型,它使用对物种有意识的嵌入来区分不同物种的DNA序列. 它在识别物种方面表现出色,即使数据有限,也提高了聚类和分类的准确性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 从基因组序列中区分物种至关重要,但具有挑战性,特别是对于缺乏参考基因组的未表征物种而言.
- 在没有参考基因组的情况下,基于嵌入的无监督方法对于物种识别至关重要.
- 现有的基因组基础模型需要改进,以在易出错的长时间读取的DNA序列上提供强大的性能.
研究的目的:
- 引入DNABERT-S,这是一个定制的基因组模型,用于使用DNA序列进行无监督物种差异化.
- 开发物种意识的嵌入,有效地集群和分离物种的DNA序列.
- 在标签稀缺和低数据场景中改进模型的性能.
主要方法:
- DNABERT-S建立在DNABERT-2基础模型的基础上.
- 介绍了多重实例混合 (MI-Mix),这是一个对比的学习目标,用于增强DNA序列的隐藏表示.
- 纳入课程对比学习 (C2LR) 策略,以进一步完善该模型的学习过程.
主要成果:
- 在23个不同的数据集中,DNABERT-S表现出显著的有效性,特别是在标签稀缺的条件下.
- 该模型成功地从未标记的基因组混合物中识别出两倍多的物种.
- 实现了物种聚类的调整后兰德指数 (ARI) 的翻倍,与10次射击基线相比,实现了优异的2次射击分类性能.
结论:
- DNABERT-S为从基因组序列中进行无监督物种识别提供了一个强大的工具,特别是当参考数据有限时.
- 新的MI-Mix和C2LR策略增强了该模型为DNA序列生成歧视性嵌入的能力.
- 该模型为生物多样性监测和涉及未表征物种的基因组研究提供了一个有希望的解决方案.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
The Evidence for Evolution
42.7K
Genetic variations accumulating within populations over generations give rise to biological evolution. Evolutionary changes can result in the formation of novel varieties and entire new species. These changes are responsible for the diverse forms of life inhabiting the planet. The evidence for evolution suggests that all living organisms descended from common ancestors.
42.7K
DNA as a Genetic Template
6.8K
6.8K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
DNA Isolation
193.0K
DNA from cells is required for many biotechnology and research applications, such as molecular cloning. To remove and purify DNA from cells, researchers use various methods of DNA extraction. While the specifics of different protocols may vary, some general concepts underlie the process of DNA extraction.
193.0K
Formation of Species
39.3K
Speciation describes the formation of one or more new species from one or sometimes multiple original species. The resulting species are discrete from the parent species, and barriers to reproduction will typically exist. There are two primary mechanisms, speciation with and without geographic isolation—allopatric and sympatric speciation, respectively.
39.3K

