SemiBin2:自主监督的对比学习导致了更好的MAG,用于短读和长读序列
Shaojun Pan1,2, Xing-Ming Zhao1,2,3,4, Luis Pedro Coelho1,2
1Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University, Shanghai 200433, China.
Bioinformatics (Oxford, England)
|June 30, 2023
概括
SemiBin2 通过自主监督学习改进了元基因组组装基因组 (MAG) 重建,优于以前的方法. 这种方法提高了分类效率,并降低了短读和长读测序数据的计算成本.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 从环境DNA重建基因组的过程中,元基因组对接至关重要.
- 像SemiBin这样的现有方法需要计算密集的对应注释.
- 需要更高效,更准确的元基因组分类工具.
研究的目的:
- 开发一种改进的元基因组结合方法,SemiBin2,利用自我监督学习.
- 评估SemiBin2的性能与模拟和真实数据集上的最先进的捆绑工具相比.
- 通过集成聚类算法来适应SemiBin2用于长读测序数据.
主要方法:
- 实施了一种自我监督的学习方法,从内置数据中提取特征嵌入.
- 开发了一种基于集群的DBSCAN集群算法,用于长读数据.
- 将SemiBin2与现有的元基因组分类软件进行比较.
主要成果:
- 在重建高质量的MAG中,SemiBin2显著超过SemiBin1和其他最先进的内存器.
- 与SemiBin1.1相比,SemiBin2实现了8.321.5%的高质量垃圾箱,使用更少的计算资源.
- 适应长时间读取数据的SemiBin2产生了比下一个最佳方法多13.126.3%的高质量基因组.
结论:
- SemiBin2 提供了一种更有效,更准确的解决方案,用于使用自主监督学习进行元基因组组合.
- 开发的集合集群方法将SemiBin2的实用性扩展到长读序列数据.
- 在复杂的环境样本中重建MAG的过程中,SemiBin2代表了一项重大进展.
相关概念视频
Maxam-Gilbert Sequencing
11.3K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.3K
Next-generation Sequencing
91.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
91.6K
RNA-seq
10.1K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.1K
Sanger Sequencing
755.1K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
755.1K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K


