Enhancing influenza surveillance in Australia using Google search data combined with machine learning
Aminath Shausan1, Adam Dunn2, Fiona May3
1Australian e-Health Research Centre CSIRO, Brisbane, Queensland, Australia . aminath.shausan@csiro.au.
Background:
To effectively respond to acute respiratory infections such as influenza, timely information about the spread of disease is required. Surveillance methods that train machine learning models on Google Trends data have been developed to forecast outbreaks. Our aim was to examine the feasibility of forecasting influenza rates for states and territories in Australia using Google Trends data.
Methods:
Weekly search volume data across 2018 and 2019 from Google Trends, for each state and territory in Australia excepting the Australian Capital Territory, were compared to weekly influenza notifications from the National Notifiable Disease Surveillance System (NNDSS). Four supervised machine learning regression models were developed: elastic net; support vector regression; random forest; and feedforward neural network. Models were fitted independently to data from each state and territory, considering nowcast (current week) and forecast (one and two weeks ahead) predictions.
Results:
The results show that Google search volumes are correlated with reported influenza rates over time. Random forest and elastic net models generally showed better predictive performance compared to other models. Each modelled state and territory, except the Northern Territory and Tasmania, showed at least two search queries with moderate to strong Pearson correlation with the influenza notifications. The forecast accuracy (measured by the Pearson correlation coefficient) for the best-performing model varied by state or territory (from -0.353 to 0.977) and accuracy was lower for forecasts for one week and two weeks ahead.
Conclusion:
Combinations of Google search data and machine learning showed predictive utility for forecasting weekly influenza rates in Australian jurisdictions, with lower accuracy in jurisdictions with smaller populations. Future work should investigate specific keywords by location exploring disease prediction at a more geographically specific level, with considerations for smaller populations and robustness over time.


