Related Experiment Video
Updated: Nov 7, 2025

A Component-resolved Diagnostic Approach for a Study on Grass Pollen Allergens in Chinese Southerners with Allergic Rhinitis and/or Asthma
Published on: June 4, 2017
Development of a Random Forest model for forecasting allergenic pollen in North America
Fiona Lo1, Cecilia M Bitz2, Jeremy J Hess3
1Department of Environmental and Occupational Health Sciences, School of Public Health, University of Washington, United States of America.
Abstract:
Pollen allergies have negative impacts on health. Information about airborne pollen concentration can improve symptom management by guiding choices affecting timing of medicines and pollen exposure. Observations provide accurate pollen concentrations at point locations. However, in the contiguous United States and southern Canada (CUSSC), observations are sparse, and sampling is often seasonal, intermittent or both. Modeling pollen concentration can fill in the gaps with estimates where direct observations are unavailable and also provide much-needed forecasts. The goal of this study is to develop and evaluate statistical models that predict daily pollen concentrations using a machine learning Random Forest algorithm. To evaluate our methods, we made retrospective forecasts of four pollen types (Quercus, Cupressaceae, Ambrosia and Poaceae), each in one of four CUSSC locations. Meteorological and vegetation conditions were input to the models at city and regional scales. A data augmentation technique was investigated and found to improve model skill. Models were also developed to forecast pollen in locations where there are no observations. Forecast skill in these models were found to be greater than in previous models. Nevertheless, the skill is limited by the spatiotemporal resolution of the pollen observations.
Related Concept Videos
Allergic Reactions
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

