Related Experiment Video
Updated: Jan 27, 2026

Rescue and Characterization of Recombinant Virus from a New World Zika Virus Infectious Clone
Published on: June 7, 2017
Confronting data sparsity to identify potential sources of Zika virus spillover infection among primates
Barbara A Han1, Subhabrata Majumdar2, Flavio P Calmon3
1Cary Institute of Ecosystem Studies, Box AB Millbrook, NY 12545, USA.
Abstract:
The recent Zika virus (ZIKV) epidemic in the Americas ranks among the largest outbreaks in modern times. Like other mosquito-borne flaviviruses, ZIKV circulates in sylvatic cycles among primates that can serve as reservoirs of spillover infection to humans. Identifying sylvatic reservoirs is critical to mitigating spillover risk, but relevant surveillance and biological data remain limited for this and most other zoonoses. We confronted this data sparsity by combining a machine learning method, Bayesian multi-label learning, with a multiple imputation method on primate traits. The resulting models distinguished flavivirus-positive primates with 82% accuracy and suggest that species posing the greatest spillover risk are also among the best adapted to human habitations. Given pervasive data sparsity describing animal hosts, and the virtual guarantee of data sparsity in scenarios involving novel or emerging zoonoses, we show that computational methods can be useful in extracting actionable inference from available data to support improved epidemiological response and prevention.
Insights
Identifying animal reservoirs for diseases like Zika virus (ZIKV) is crucial. Machine learning models successfully pinpointed high-risk primate species, aiding in preventing future zoonotic spillover events.
Area of Science:
- Epidemiology
- Machine Learning
- Zoonotic Disease Surveillance
Background:
- The Zika virus (ZIKV) epidemic highlighted the threat of mosquito-borne flaviviruses.
- Sylvatic cycles in primates pose spillover risks to humans, but data on reservoirs is scarce.
- Limited surveillance data complicates identifying and mitigating zoonotic disease threats.
Purpose of the Study:
- To address data sparsity in identifying zoonotic reservoirs.
- To develop computational methods for predicting high-risk animal species.
- To improve epidemiological response and prevention strategies for emerging zoonoses.
Main Methods:
- Employed Bayesian multi-label learning, a machine learning technique.
- Utilized multiple imputation to handle missing primate trait data.
- Developed predictive models for flavivirus-positive primates.
Main Results:
- Models achieved 82% accuracy in distinguishing flavivirus-positive primates.
- Identified primate species with high spillover risk are often adapted to human environments.
- Demonstrated the utility of computational methods in data-scarce zoonotic scenarios.
Conclusions:
- Machine learning effectively extracts actionable insights from limited data.
- Predictive modeling can guide surveillance and prevention of zoonotic diseases.
- Understanding host adaptation is key to mitigating spillover risk.
Related Concept Videos
GIS Software, Hardware, and Sources of GIS Data
What are Viruses?
Potential Energy
Chemical bonds that form attractive forces between atoms also contain potential energy, called chemical energy. When a chemical reaction...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Sinusoidal Sources
In homes, the power supplies use sinusoidal sources to provide electricity. These sources generate a voltage that varies sinusoidally...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...

