以Unitig为中心的全基因组机器学习方法用于预测抗生素耐药性和发现细菌菌株中的新型耐药性基因
Duyen Thi Do1, Ming-Ren Yang1,2, Tran Nam Son Vo3
1Graduate Institute of Biomedical Informatics, College of Medical Science and Technology, Taipei Medical University, Taipei, Taiwan.
概括
这项研究引入了一种新的机器学习方法,使用泛基因组分析来预测抗菌素耐药性 (AMR). 该方法准确地识别了细菌中的抗生素耐药性,并发现了新的耐药性基因,改善了已知基因数据库之外的AMR预测.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 目前的抗菌素耐药性 (AMR) 预测方法依赖于已知的基因或参考基因组,由于不完全覆盖耐药性机制和遗传多样性,限制了准确性.
- 现有的基因组方法在准确预测AMR方面面临挑战,原因是大量的遗传变异和尚未发现的耐药性决定因素.
研究的目的:
- 开发和验证基于泛基因组的机器学习方法,以准确预测抗菌素耐药性 (AMR).
- 通过先进的基因组分析,识别新型AMR基因并扩大已知的抗药性机制库.
主要方法:
- 从数千个微生物基因组中构建了紧的德布鲁金图 (cDBGs).
- 收集了独特序列 (unitigs) 的存在/缺失模式,用于特征提取.
- 应用基于特征选择的机器学习模型用于AMR分类.
主要成果:
- 使用单元中心泛基因组的机器学习模型在预测抗生素耐药性或易感性方面表现出显著的前景.
- 在训练数据集上,接收器操作特征曲线 (AUC) 下的面积>0.929,在独立验证数据集上约为0.77,实现了高预测性能.
- 鉴定了以前未知的抗药性基因,扩大了已知的AMR基因库,并提供了对新型抗药性机制的见解.
结论:
- 拟议的基于unitig的泛基因组特征集有效地使得针对AMR病原体构建精确的机器学习预测器.
- 这种方法为扩大已知的AMR基因数据库和产生关于细菌AMR机制的新假设提供了有价值的见解.
相关概念视频
Antibiotic Selection
53.2K
Overview
53.2K
Genomic DNA in Prokaryotes
43.8K
The genome of most prokaryotic organisms consists of double-stranded DNA organized into one circular chromosome in a region of cytoplasm called the nucleoid. The chromosome is tightly wound, or supercoiled, for efficient storage. Prokaryotes also contain other circular pieces of DNA called plasmids. These plasmids are smaller than the chromosome and often carry genes that confer adaptive functions, such as antibiotic resistance.
Genomic Diversity in Bacteria
Although bacterial genomes are much...
Genomic Diversity in Bacteria
Although bacterial genomes are much...
43.8K
Genome Size and the Evolution of New Genes
7.9K
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
7.9K


