Related Experiment Video
Updated: Jan 30, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Predicting ecological risks of heavy metals in watersheds based on interpretable machine learning models: under the
Hong Chen1,2, Meiling Kong2, Zhenghua Wu2
1School of Environment and Energy Engineering, Anhui Jianzhu University, Hefei, 230601, China.
Synthetic data generation using SMOGN significantly improved machine learning model accuracy for predicting heavy metal toxicity in watershed ecosystems. This approach enhances ecological risk assessment and supports pollution control strategies.
Area of Science:
- Environmental Science
- Ecotoxicology
- Data Science
Background:
- Heavy metal contamination threatens watershed ecosystems, necessitating accurate ecological risk prediction.
- Limited ecotoxicological data due to high costs hinders machine learning applications.
Purpose of the Study:
- To address data scarcity in ecotoxicology by augmenting toxicity datasets using synthetic data.
- To develop and compare machine learning models for predicting heavy metal toxicity (Cr, Mn, Cu).
- To interpret key toxicity drivers and assess ecological risks in the Chaohu Lake Basin.
Main Methods:
- Employed Synthetic Minority Oversampling Technique for Regression with Gaussian Noise (SMOGN) to generate synthetic ecotoxicity data.
- Developed and compared Random Forest (RF), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM) regression models.
- Utilized Shapley Additive Explanations (SHAP) for model interpretability and Risk Quotient (RQ) for ecological risk assessment.
Main Results:
- Synthetic data augmentation substantially improved model performance, with RF R² increasing from 0.696 to 0.977.
- All tested models (RF, XGBoost, SVM) demonstrated higher accuracy with augmented data compared to the original dataset.
- Ecological risk assessment revealed elevated risks for Chromium (Cr) and Copper (Cu) in the Chaohu Lake Basin.
Conclusions:
- SMOGN is a feasible and effective method for overcoming data scarcity in ecotoxicology.
- Augmented datasets enhance the accuracy of machine learning models for predicting heavy metal toxicity.
- The study provides crucial data and evidence for managing heavy metal contamination in the Chaohu Lake Basin.
More Related Videos
Related Concept Videos
Ecological Disturbance
Bioequivalence Data: Statistical Interpretation
Ecological Succession
Ecological Niches
Bonding in Metals
Simplified Synchronous Machine Model
In this model, each generator is connected to a...

