一个数据驱动的框架,以确定影响土壤重金属污染的因素,使用随机森林和双变的本地莫兰的I:一个案例研究
Rui Zhou1, Jian Chen2, Shiwen Cui3
1College of New Energy and Environment, Jilin University, Changchun, 130012, China.
Journal of environmental management
|January 22, 2025
概括
一个新的数据驱动框架确定了影响重金属 (HM) 土壤污染的关键因素. 这种方法有助于确定污染源,并指导工业区的有针对性的预防和控制措施.
科学领域:
- 环境科学 环境科学
- 地质化学 地质化学
- 数据科学数据科学数据科学
背景情况:
- 土壤中的重金属 (HM) 污染对环境构成重大风险.
- 对HM污染的可追溯性分析通常受到影响因素及其空间相关性数据不足的限制.
研究的目的:
- 建立一个新的数据驱动框架,以识别和量化影响土壤HM度的因素.
- 在一个工业化地区分析HM度和环境共变量之间的空间相关性.
主要方法:
- 利用了来自中国广东省的577个土壤样本和18个环境共变量的数据集.
- 雇佣随机森林 (RF) 用于对影响因素的定量贡献分析.
- 应用了多变局部莫兰的I (BLMI) 来生成空间聚类地图.
主要成果:
- 确定了特定HM的关键因素:Cd (加油站,铁路),As (地下水深度,海拔),Pb (土壤pH,危险废物场所) 和Cr (矿山尾矿,降雨).
- 对于Cd,As,Pb和Cr度的18个因素的定量贡献.
- 生成空间聚类地图,显示在研究区域中间高HM度和复杂的人类活动.
结论:
- 数据驱动的框架有效地识别和量化影响土壤HM污染的因素.
- 由于高度和复杂的人类活动,研究区域的中部地区需要优先关注预防和控制措施.
- 该框架提供了全面的信息,以了解和管理HM污染.
相关概念视频
Strategies for Assessing and Addressing Confounding
82
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
82
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K


