基因拷贝数特征比SNP更好地泛化,用于预测Staphylococcus aureus中的抗菌素耐药性
Bruna F Fistarol1, Joao D Gervasio2, Gergely J Szöllősi3,4
1Model-Based Evolutionary Genomics Unit, Okinawa Institute of Science and Technology, Okinawa, Japan. brunaffistarol@gmail.com.
npj antimicrobials and resistance
|December 16, 2025
概括
使用细菌基因组序列预测抗菌耐药性 (AMR) 是至关重要的. 全基因组基因拷贝数模型显著优于单核酸多态 (SNP) 模型,特别是在新型细菌系中.
科学领域:
- 基因组学就是基因组学.
- 微生物学 微生物学
- 计算生物学 计算生物学
背景情况:
- 从基因组序列来准确预测抗菌素耐药性 (AMR) 对有效治疗至关重要.
- 目前使用精选标记面板或核心基因组单核酸多态 (SNP) 的方法,难以将其推广到新的细菌系.
研究的目的:
- 评估泛基因组基因拷贝数特征的有效性,以预测黄金葡萄球菌*中的AMR.
- 将基因内容模型与SNP模型的性能进行比较,特别是在涉及新型细菌系的场景中.
主要方法:
- 在泛基因组特征上利用渐变增强的决策树集 (XGBoost),编码同源基因拷贝数 (包括缺席).
- 基因含量模型与基于SNP的模型在6种抗生素和4255种*金黄色葡萄球菌*隔离物中进行了比较.
- 进行了谱系的评估,以评估模型对未见的细菌群的概括性.
主要成果:
- 基因拷贝数模型实现了高宏平均F1得分 (0.925-0.988),超过了基于SNP的模型 (0.838-0.935).
- 基因含量模型在基因系外评估下表现优异 (F1=0.875和0.904),而SNP模型显著退化 (F1=0.557和0.638).
- 特征切除揭示了AMR预测信号分布在众多基因家族中,支持更广泛的跨谱系概括.
结论:
- 全基因组基因拷贝数表示为SNP-only方法提供了强大的替代方案,用于AMR预测.
- 这种基于基因内容的方法增强了基于基因组的AMR预测,使其适用于临床和流行病学数据集,即使覆盖率低的测序.
- 这些发现强调了包括复制数变异在内的全面基因内容对于开发可概括的AMR预测模型的重要性.
相关概念视频
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Antibiotic Selection
59.3K
Overview
59.3K
Single Nucleotide Polymorphisms-SNPs
17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K
Genome Size and the Evolution of New Genes
8.9K
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
8.9K
Modern Molecular Taxonomy
543
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
543
Genome Copying Errors
5.0K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
5.0K


