Related Experiment Video
Updated: May 28, 2026

Use of Principal Components for Scaling Up Topographic Models to Map Soil Redistribution and Soil Organic Carbon
Published on: October 16, 2018
Machine learning for the spatial prediction of soil and groundwater contamination: An integrative global review
He Wu1, Zhuo Zhang2, Panpan Wang3
1School of Land Science and Technology, China University of Geosciences (Beijing), Beijing 100083, China.
Abstract:
The spatial distribution of soil and groundwater pollutants is critical for effective remediation. Machine learning methods are increasingly applied in predicting pollutant distributions due to their efficiency, low cost, and ability to capture complex nonlinear relationships. To date, however, systematic and quantitative reviews of the full modeling workflow remain limited. In this study, 214 publications worldwide were systematically reviewed and quantitatively analyzed across three core dimensions: data, models, and driving factors. The analysis revealed that, although data types and sources are diverse, Challenges remain in obtaining high-quality data and effectively integrating heterogeneous data sources. Most studies partitioned datasets into training and testing subsets, with training proportions typically ranging from 60% to 80%, and applied evaluation metrics tailored to research objectives. In terms of models, traditional approaches are well-established and generally perform satisfactorily, whereas complex and emerging methods, such as graph neural networks, offer higher predictive accuracy and broader spatial coverage, and warrant further exploration with a focus on enhancing interpretability. The selection of driving factors is primarily influenced by pollutant characteristics and the intrinsic properties of the soil or groundwater, as well as the surrounding environmental context, which collectively determine the relative importance and contribution of different variables in prediction models. Collectively, this review provides methodological guidance for feature design, model choice, and the application of interpretability tools in future pollution prediction studies.