Related Experiment Video
Updated: Aug 27, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
A Machine Learning-Based Water Potability Prediction Model by Using Synthetic Minority Oversampling Technique and
Jinal Patel1, Charmi Amipara1, Tariq Ahamed Ahanger2
1Department of Computer Science and Engineering Pandit Deendayal Energy University, Gandhinagar, Gujarat, India.
Machine learning models can predict water quality. Random Forest and Gradient Boost achieved 81% accuracy, with explainable AI identifying key predictive features for better transparency.
Area of Science:
- Environmental Science
- Computer Science
- Data Science
Background:
- Water quality degradation is a significant global issue.
- Accurate water quality prediction is crucial for environmental management.
- Existing machine learning models lack transparency.
Purpose of the Study:
- To compare machine learning algorithms for water quality classification.
- To enhance model transparency using explainable AI (XAI).
- To identify key features influencing water quality predictions.
Main Methods:
- Comparative analysis of Support Vector Machine (SVM), Decision Tree (DT), Random Forest, Gradient Boost, and Ada Boost algorithms.
- Dataset normalization using Z-score.
- Handling imbalanced data with Synthetic Minority Oversampling Technique (SMOTE).
- Feature importance analysis using Local Interpretable Model-agnostic Explanations (LIME).
Main Results:
- Random Forest and Gradient Boost models achieved the highest accuracy at 81%.
- Explainable AI (XAI) techniques, specifically LIME, were successfully applied.
- Key features contributing to water quality classification were identified.
Conclusions:
- Machine learning, particularly Random Forest and Gradient Boost, offers effective water quality classification.
- Integrating XAI improves model interpretability, addressing a critical limitation.
- Feature importance analysis provides insights for targeted water quality management strategies.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Steps in Outbreak Investigation
Non-equilibrium in the Cell
Mechanistic Models: Compartment Models in Individual and Population Analysis
Modeling and Similitude
Testing Water Quality

