使用可解释的机器学习模型了解地下水干旱的进化过程和原因
Zhiyuan Gan1,2, Xianjun Xie1,2, Chunli Su3,4
1School of Environmental Studies, China University of Geosciences, Wuhan, 430074, China.
Scientific reports
|July 2, 2025
概括
使用机器学习和SHAP分析改进了地下水干旱预测,确定长期气象干旱是关键驱动因素. 未来气候情景预计干旱的严重程度和范围会增加.
科学领域:
- 水文学的水文学
- 气候科学 气候科学
- 机器学习 机器学习
背景情况:
- 由于有限的直接观测,对地下水干旱的评估具有挑战性.
- 了解地下水干旱演变对于水资源管理至关重要.
研究的目的:
- 开发和评估用于预测地下水干旱的机器学习模型.
- 通过SHAP分析确定影响地下水干旱的关键因素.
- 在气候变化情景下预测未来的地下水干旱趋势.
主要方法:
- 采用机器学习模型,包括由Sparrow搜索算法 (SSA) 优化的XGBoost.
- 使用Shapley添加式解释 (SHAP) 进行模型解释性和特征重要性分析.
- 在西河平原 (WLRP) 评估了八种地下水干旱预测模型.
主要成果:
- 经过SSA优化的XGBoost模型表现出高性能 (AUC:0.922,F1得分:0.84).
- 在12个月和24个月的标准化降水蒸发指数 (SPEI) 被确定为关键预测指标.
- 长期的气象干旱,过度开采和城市化被发现加剧了地下水干旱.
结论:
- 机器学习和SHAP分析为了解地下水干旱提供了一个强大的框架.
- 长期的气象干旱显著影响地下水干旱,与其他因素的相互作用.
- 预计未来的气候变化 (SSP5-8.5) 将增加地下水干旱的频率,范围和严重程度.
相关概念视频
Steps in Outbreak Investigation
213
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
213
Responses to Drought and Flooding
11.0K
Water plays a significant role in the life cycle of plants. However, insufficient or excess of water can be detrimental and pose a serious threat to plants.
11.0K
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Precipitation Processes
608
The experimental conditions in a gravimetric analysis should be optimized to maximize the particle size and purity of the obtained precipitate. Ideally, the concentration of the precipitating reagent should be low with effective stirring to maintain low relative supersaturation for the growth of large crystals. In homogeneous precipitation, the precipitant is slowly generated by a chemical reaction in the solution to avoid local reagent excesses. For example, urea decomposes gradually to...
608
Precipitation Gravimetry
7.6K
Precipitation gravimetry is based on converting an analyte into a sparingly soluble precipitate, which is separated by filtration and weighed. An ideal precipitate should be pure, insoluble, of known composition, and easily filtered from the reaction mixture.
In determining nickel by gravimetric analysis, a precipitant of ethanolic dimethylglyoxime is added to a hot nickel salt solution. This is quickly followed by the dropwise addition of dilute ammonia solution until precipitation occurs. A...
In determining nickel by gravimetric analysis, a precipitant of ethanolic dimethylglyoxime is added to a hot nickel salt solution. This is quickly followed by the dropwise addition of dilute ammonia solution until precipitation occurs. A...
7.6K
Mechanistic Models: Compartment Models in Individual and Population Analysis
89
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
89


