Related Experiment Video
Updated: Oct 25, 2025

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
Published on: January 13, 2023
Machine learning methods for imbalanced data set for prediction of faecal contamination in beach waters
Mathias Bourel1, Angel M Segura2, Carolina Crisci2
1IMERL, Facultad de Ingeniería, Universidad de la República, Montevideo, Uruguay; Departamento de Modelización Estadística de Datos e Inteligencia Artificial (MEDIA), Centro Universitario Regional Este, Universidad de la República, Rocha, Uruguay.
Predicting rare water contamination events at recreational beaches is challenging due to imbalanced data. Machine learning models, particularly stratified Random Forest, show promise in improving prediction accuracy for fecal coliforms.
Area of Science:
- Environmental Science
- Public Health
- Data Science
Background:
- Statistical models aid in managing health risks at recreational beaches.
- Predicting extreme water contamination events is difficult due to imbalanced datasets where rare events are infrequent.
Purpose of the Study:
- To evaluate machine learning techniques and metrics for modeling imbalanced data in predicting water contamination.
- To identify optimal methods for anticipating water contamination and mitigating health risks.
Main Methods:
- Utilized simulated datasets and a real-world database of fecal coliform abundance from 21 Uruguayan beaches over 10 years.
- Tested various machine learning algorithms and data pre-treatment methods like upsampling and SMOTE.
- Assessed model performance using True Positive Rate (TPR) and False Positive Rate (FPR), not just accuracy.
Main Results:
- Most machine learning techniques require data pre-treatment (e.g., upsampling) to enhance performance on imbalanced datasets.
- Stratified Random Forest demonstrated superior performance, improving TPR by 50% compared to baseline.
- Support Vector Machines with SMOTE and Adaboost with SMOTE also yielded strong results.
Conclusions:
- Machine learning models are sensitive to data imbalance, necessitating specific pre-processing techniques.
- TPR and FPR are more reliable metrics than accuracy for evaluating models with imbalanced data.
- Combining modeling strategies is crucial for improving the prediction of water contamination events and safeguarding public health.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Steps in Outbreak Investigation
Mechanistic Models: Compartment Models in Individual and Population Analysis
Quantifying and Rejecting Outliers: The Grubbs Test

