Related Experiment Video
Updated: Mar 29, 2026

10:26
Author Spotlight: A Cost-Effective Genomic Workflow for Advancing Rabies Control in Resource-Limited Settings
Published on: August 18, 2023
6.7K
An improved dataset for predicting mammal infecting viruses from genetic sequence information
Tyler Reddy1, Austin Schneider2, Aaron R Hall1
1CAI-1: Applied Computer Science, Los Alamos National Laboratory, Los Alamos, New Mexico.
Plos Computational Biology
|March 27, 2026
Summary
Developing machine learning (ML) models to predict virus host infections is challenging. A standardized dataset and improved methods show better prediction of mammal-infecting viruses, highlighting the importance of taxonomic rank and reduced phylogenetic distance.
Area of Science:
- Computational Biology and Bioinformatics
- Virology
- Machine Learning
Background:
- Machine learning (ML) models for identifying human-infecting viruses from genomic sequences have shown varying success.
- Direct comparison of these models is difficult due to differing datasets, evaluation metrics, and data splitting strategies.
- Previous work by Mollentze et al. provided a foundation for host-virus record curation.
Purpose of the Study:
- To present a standardized, expanded dataset of mammal-infecting and non-infecting viral pathogens.
- To evaluate the performance of eight ML models on this dataset for predicting mammal-infecting viruses.
- To enable standardized comparisons of ML methods for predicting human host infections.
Main Methods:
- Refined and expanded a dataset of mammal-infecting and non-infecting viral pathogens, incorporating new host labels (primate and mammal).
- Evaluated eight ML models on the standardized dataset using ROC AUC (Receiver Operating Characteristic Area Under the Curve) as a performance metric.
- Compared model performance using random data splitting versus original assignments and analyzed the impact of phylogenetic distance and viral family overlap.
Main Results:
- Random data splitting in the improved dataset increased the average ROC AUC for predicting human infection from 0.663 to 0.784.
- Prediction of mammal infection was most reliable at the broadest host category, achieving an ROC AUC of 0.850.
- Models performed no better than random chance (ROC AUC 0.50) when training and test sets had no overlap in viral families, indicating challenges in out-of-sample prediction.
Conclusions:
- Virus host infection classification is more tractable at higher taxonomic ranks.
- Reducing phylogenetic distance between training and test sets significantly improves predictive performance.
- Peptide k-mer features may negatively impact out-of-sample model performance, and predicting virus host jumps remains a significant challenge.
Related Concept Videos
Viral Mutations
40.7K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
40.7K
Viral Recombination
25.6K
Cells are sometimes infected by more than one virus at once. When two viruses disassemble to expose their genomes for replication in the same cell, similar regions of their genomes can pair together and exchange sequences in a process called recombination. Alternatively, viruses with segmented genomes can swap segments in a process called reassortment.
25.6K

