Predictive modeling of Pseudomonas syringae virulence on bean using gradient boosted decision trees
Renan N D Almeida1, Michael Greenberg1, Cedoljub Bundalovic-Torma1
1Department of Cell & Systems Biology, University of Toronto, Toronto, Canada.
Pseudomonas syringae pathovar phaseolicola (PG3) strains show higher host specificity on common bean than pathovar syringae (PG2) strains. Machine learning accurately predicts bacterial virulence using whole genome data, demonstrating potential for understanding host adaptation.
Area of Science:
- Bacteriology
- Plant Pathology
- Genomics
- Machine Learning
Background:
- Pseudomonas syringae pathovar designations are assumed to reflect host specificity, but this is rarely tested.
- Understanding host specificity in P. syringae is crucial for managing crop diseases.
Purpose of the Study:
- To assess the virulence of diverse P. syringae isolates on common bean (Phaseolus vulgaris).
- To determine host specificity differences between P. syringae phylogroup 2 (PG2) and phylogroup 3 (PG3) bean isolates.
- To develop a machine learning model for predicting P. syringae virulence on bean.
Main Methods:
- A rapid seed infection assay was used to measure virulence of 121 P. syringae isolates on common bean.
- Gradient boosting machine learning models were trained using whole genome k-mers, type III secreted effector k-mers, and effector/phytotoxin presence/absence data.
- Model predictions were functionally validated on 16 strains.
Main Results:
- Bean-infecting P. syringae isolates were generally more virulent on bean than non-bean isolates.
- PG3 bean isolates were significantly more virulent on bean than PG3 non-bean isolates, indicating higher host specificity.
- Machine learning models using whole genome data achieved high accuracy in predicting virulence (mean absolute error = 0.05).
- 94% of predicted virulence values for 16 strains fell within validated bounds.
Conclusions:
- PG3 P. syringae strains exhibit greater host specificity to common bean compared to PG2 strains.
- P. syringae PG2 strains may have evolved a broader host range, reflecting a different lifestyle.
- Machine learning is a powerful tool for predicting bacterial virulence and understanding host-specific adaptation in plant pathogens.
More Related Videos
12:26Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
Published on: October 11, 2016
09:21Author Spotlight: Generating Neuronal Phenotypic Profiles - A Protocol to Culture and Image Human Midbrain Dopaminergic Neurons
Published on: July 7, 2023
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Light Acquisition
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
