Related Experiment Video
Updated: Jun 3, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Identifying predictors of spatiotemporal variations in residential radon concentrations across North Carolina using
Zhenchun Yang1, Lauren Prox2, Clare Meernik3
1Duke Global Health Institute, Durham, NC, 27708, United States.
Abstract:
Radon is a naturally occurring radioactive gas derived from the decay of uranium in the Earth's crust. Radon exposure is the leading cause of lung cancer among non-smokers in the US. Radon infiltrates homes through soil and building foundations. This study advances methodologies for assessing residential radon exposure by leveraging a comprehensive dataset of 126,382 short-term (2-7 days) radon test results collected across North Carolina from 2010 to 2020. Employing a combination of linear regression and advanced machine learning techniques, including random forest models. Analysis through linear regression, linear mixed-effects models (LME), and generalized additive models (GAM) using the first-time tested radon levels reveals that elevation, proximity to geological faults, and soil moisture are pivotal in determining radon concentration. Specifically, elevation consistently shows a positive relationship with radon levels across models (linear regression: β = 0.12, p < 0.001; LME: β = 0.17, p < 0.001; GAM: β = 0.11, p < 0.001). Conversely, the distance to geological faults negatively correlates with radon concentration (linear regression: β = -0.11, p < 0.001; LME: β = -0.06, p < 0.001; GAM: β = -0.07, p < 0.001), indicating lower radon levels further from faults. Using the random forest model, our study identifies the most influential environmental predictors of first-time tested radon levels. Elevation is the most influential variable, followed by median instantaneous surface pressure and soil moisture in the upper 10 cm layer, illustrating the significant role of geological and immediate surface conditions. Additional important factors include precipitation, mean temperature, and deeper soil moisture levels (40-200 cm), which underscores the influence of climate on radon variability. Root zone soil moisture and the Normalized Difference Vegetation Index (NDVI) also contribute to predicting radon levels, reflecting the importance of soil and vegetation dynamics in radon emanation. By integrating multiple statistical models, this research provides a nuanced understanding of the predictors of radon concentration, enhancing predictive accuracy and reliability.
More Related Videos
14:27Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Steps in Outbreak Investigation
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as: