Related Experiment Video
Updated: Jan 17, 2026

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
The comparison between multiple linear regression and random forest model in predicting environmental noise and its
Chui Hei Wong1, Zhiyuan Li2, Steve Hung Lam Yim3
1Jockey Club School of Public Health and Primary Care, The Chinese University of Hong Kong, Hong Kong Special Administrative Region.
Abstract:
Environmental noise exposure is suggested to be linked with various chronic diseases' development, including cardiovascular diseases and cognitive decline. However, traditional acoustic models are not flexible enough in estimating participants' noise exposure with diverse exposure sources in epidemiological studies. The Land-use regression (LUR) model has been recommended as a better alternative in predicting noise levels with less data input. The comparison of multiple linear regression (MLR) and machine learning algorithms in developing the LUR model is required to understand how machine learning algorithms overcome the non-linearity issue of predictor variables that the MLR model cannot solve. Random forest is the more favorable machine learning algorithm in developing LUR model due to its ability in handling outliers and overfitting. Summer and winter measurements for A-weighted equivalent sound pressure levels over 24 h (Leq,24h) and nighttime (Lnight) with the frequency components were conducted at 102 sites during 2019-2020. The noise parameters together with various traffic, population and land-use variables were used to estimate the spatial variability of environmental noise in Hong Kong. Random forest (RF) models performed better than the MLR models with more predictors and in higher leave-one-out-cross-validation R2 in Leq,24h (RF: 0.79, MLR: 0.70). MLR model attained a better performance than RF in fitted R2 for all night noise parameters. Additionally, MLR outperformed RF in cross-validation mean absolute error and root-mean-square error in some noise frequencies at night. Vegetation, industrial area and buses were mostly included in both models of high importance. Our findings demonstrate that the RF model can generate credible predictions comparable to those of the MLR model for future epidemiological studies focusing on noise-related health outcomes.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
04:35Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...
Expected Frequencies in Goodness-of-Fit Tests
Random Error