4CAC:使用机器学习和装配图的元基因组结合物的4类分类器
Lianrong Pu1,2, Ron Shamir1
1The Blavatnik School of Computer Science, Tel Aviv University, Tel Aviv, Israel.
Nucleic acids research
|September 17, 2024
概括
一个名为4CAC的新工具准确地识别了微生物群落中的病毒,等离子体和微核细胞. 这一进步提高了对这些微小但重要的微生物参与者及其在基因转移中的作用的理解.
科学领域:
- 微生物生态学 微生物生态学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 微生物群落包含细菌,古生物,病毒,等离子体和微核细胞.
- 病毒,等离子体和微核细胞对于水平基因转移和抗生素耐药性至关重要,但由于识别挑战,它们经常被忽视.
- 现有的分类器与类不平衡作斗争,导致小微生物类的识别不佳.
研究的目的:
- 开发一个新的分类器,4CAC,用于同时识别病毒,等离子体,微核细胞和 prokaryotes 在元基因组组合.
- 在微生物社区分析中解决阶级失衡问题.
- 为了提高识别小微生物成分的准确性和效率.
主要方法:
- 开发了4CAC,一种使用序列长度调整的XGBoost模型和组装图信息进行四向分类的分类器.
- 在模拟和真实元基因组数据集上评估了4CAC.
- 将4CAC的性能与现有分类器进行比较.
主要成果:
- 在识别小微生物类别方面,4CAC显著优于现有的分类器,特别是在短阅读中.
- 4CAC在长时间读取时具有优势,除非次要类丰度非常低.
- 与其他方法相比,4CAC实现了 1-2 个数量级更快的处理速度.
结论:
- 4CAC在元基因组数据中对识别病毒,等离子体和微核细胞进行了实质性改进.
- 4CAC的速度和准确性使其成为微生物社区分析的宝贵工具.
- 4CAC增强了我们对微生物小成分在生态和进化过程中的作用的理解.
更多相关视频
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
3.9K
09:06High-throughput Identification of Gene Regulatory Sequences Using Next-generation Sequencing of Circular Chromosome Conformation Capture 4C-seq
Published on: October 5, 2018
10.2K
相关概念视频
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Classification of Systems-I
177
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
177
Aggregates Classification
306
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
306
Classification of Systems-II
137
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
137
Comparing Mitochondrial, Chloroplast, and Prokaryotic Genomes
12.0K
The present-day mitochondrial and chloroplast genomes have retained some of the characteristics of their ancestral prokaryotes and also have acquired new attributes during their evolution within eukaryotic cells. Like prokaryotic genomes, mitochondrial and chloroplast genomes neither bind with histone-like proteins nor show complex packaging into chromosome-like structures, as observed in eukaryotes. Unlike mitotic cell divisions observed in eukaryotic cells, mitochondria and chloroplasts...
12.0K
Oligosaccharide Assembly
2.8K
Protein glycosylation starts in the ER lumen and continues in the Golgi apparatus. Glycosyltransferases catalyze the addition of sugar molecules or glycosylation of proteins. Usually, these enzymes add sugars to the hydroxyl groups of selected serine or threonine residues to form O-linked glycans or the amino groups of asparagine residues to form N-linked glycans. Different positions on the same polypeptide chain can contain differently linked glycans.
Multiple sugar molecules that may or may...
Multiple sugar molecules that may or may...
2.8K
