Comparison of land use regression and random forests models on estimating noise levels in five Canadian cities
Ying Liu1, Sophie Goudreau2, Tor Oiamo3
1Canadian Urban Environmental Health Research Consortium, Canada; Department of Environmental and Occupational Health, School of Public Health, University of Montreal, Montreal, QC H3C 3J7, Canada.
Abstract:
Chronic exposure to environment noise is associated with sleep disturbance and cardiovascular diseases. Assessment of population exposed to environmental noise is limited by a lack of routine noise sampling and is critical for controlling exposure and mitigating adverse health effects. Land use regression (LUR) model is newly applied in estimating environmental exposures to noise. Machine-learning approaches offer opportunities to improve the noise estimations from LUR model. In this study, we employed random forests (RF) model to estimate environmental noise levels in five Canadian cities and compared noise estimations between RF and LUR models. A total of 729 measurements and 33 built environment-related variables were used to estimate spatial variation in environmental noise at the global (multi-city) and local (individual city) scales. Leave one out cross-validation suggested that noise estimates derived from the RF global model explained a greater proportion of variation (R2: RF = 0.58, LUR = 0.47) with lower root mean squared errors (RF = 4.44 dB(A), LUR = 4.99 dB(A)). The cross-validation also indicated the RF models had better general performance than the LUR models at the city scale. By applying the global models to estimate noise levels at the postal code level, we found noise levels were higher in Montreal and Longueuil than in other major Canadian cities.
Related Concept Videos
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Random Error
Estimating Population Standard Deviation
Expected Frequencies in Goodness-of-Fit Tests
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Distributions to Estimate Population Parameter


