Related Experiment Video
Updated: Jun 26, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Transferability of Machine Learning Models for Geogenic Contaminated Groundwaters
Hailong Cao1, Xianjun Xie2,3, Ziyi Xiao2,3
1College of Resources and Environment, Yangtze University, Wuhan 430100, China.
Transferring machine learning models for groundwater contamination is feasible when using hydrochemical data. Adding local data significantly improves model accuracy, highlighting the importance of predictor types and data informing for successful model transfer.
Area of Science:
- Environmental Science
- Hydrogeology
- Machine Learning Applications
Background:
- Machine learning (ML) models are effective for identifying geogenic contaminated groundwaters.
- Developing ML models in data-scarce regions is challenging due to the requirement for extensive training datasets.
- Model transferability offers a potential solution for applying ML in regions with limited available data.
Purpose of the Study:
- To investigate the transferability of high fluoride groundwater ML models between different basins within the Shanxi Rift System.
- To evaluate the influence of six factors: modeling methods, predictor types, data size, sample/predictor ratio (SPR), predictor range, and data informing on model transferability.
- To identify key factors that enhance or hinder the successful transfer of ML models for groundwater contamination assessment.
Main Methods:
- Exploration of ML model transferability using hydrochemical and surface parameters as predictors.
- Assessment of the impact of data informing (adding data from target regions to training sets) on transferability.
- Statistical analysis (stepwise regression) to determine significant factors influencing transferability.
- Application of the t-distributed Stochastic Neighbor Embedding (t-SNE) algorithm for data visualization and comparison across basins.
Main Results:
- Successful transferability of high fluoride groundwater models was achieved only when predictors were based on hydrochemical parameters, not surface parameters.
- Data informing, by incorporating samples from challenging regions, significantly enhanced model transferability.
- Stepwise regression confirmed hydrochemical predictors and data informing as significant positive factors, while data size, SPR, and predictor range showed no significant effect.
- Advanced ML models (random forests, artificial neural networks) did not consistently outperform logistic regression in terms of transferability.
- t-SNE analysis highlighted the critical role of predictor types in enabling effective data representation and model transfer across basins.
Conclusions:
- Hydrochemical parameters are crucial for the successful transferability of machine learning models in groundwater contamination studies.
- Data informing is a vital strategy for improving the performance of transferred models in data-limited or challenging hydrogeological settings.
- The choice of predictor type is more critical for model transferability than the complexity of the machine learning algorithm used.
Related Concept Videos
Typical Model Studies
Mechanistic Models: Compartment Models in Individual and Population Analysis
Levels of Use of a GIS
Bioremediation
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Modeling and Similitude

