Related Experiment Video
Updated: May 7, 2025

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
Performance of Conditional Random Forest and Regression Models at Predicting Human Fecal Contamination of Produce
Jessica Hofstetter1, David A Holcomb2, Amy M Kahler2
1Waterborne Disease Prevention Branch, Centers for Disease Control and Prevention, Atlanta, Georgia 30333, United States; Chenega Enterprise Systems & Solutions, LLC, Chesapeake, Virginia 23320, United States; Department of Horticulture, Auburn University, Auburn, Alabama 36849, United States.
Abstract:
Irrigating fresh produce with contaminated water contributes to the burden of foodborne illness. Identifying fecal contamination of irrigation waters and characterizing fecal sources and associated environmental factors can help inform fresh produce safety and health hazard management. Using two previously collected data sets, we developed and evaluated the performance of logistic regression and conditional random forest models for predicting general and human-specific fecal contamination of ponds in southwest Georgia used for fresh produce irrigation. Generic Escherichia coli served as a general fecal indicator, and human-associated Bacteroides (HF183), crAssphage, and F+ coliphage genogroup II were used as indicators of human fecal contamination. Increased rainfall in the previous 7 days and the presence of a building within 152 m (a proxy for proximity to septic systems) were associated with increased odds of human fecal contamination in the training data set. However, the models did not accurately predict the presence of human-associated fecal indicators in a second data set collected from nearby irrigation ponds in different years. Predictive statistical models should be used with caution to assess produce irrigation water quality as models may not reliably predict fecal contamination at other locations and times, even within the same growing region.
More Related Videos
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Survival Tree
Building a Survival Tree
Constructing a...
Receiver Operating Characteristic Plot
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...

