Predicting Groundwater PFOA Exposure Risks with Bayesian Networks: Empirical Impact of Data Preprocessing on Model
Runwei Li1,2, Jacqueline MacDonald Gibson2
1Department of Civil Engineering, New Mexico State University, 3035 S Espina St, Las Cruces, New Mexico 88003, United States.
Data preprocessing significantly impacts machine learning models predicting per- and polyfluoroalkyl substances (PFAS) exposure risks in groundwater. Despite variations, models accurately identified high-risk wells, highlighting a data quality versus model performance trade-off.
Area of Science:
- Environmental Science
- Data Science
- Toxicology
Background:
- Per- and polyfluoroalkyl substances (PFAS) are widespread environmental contaminants of growing concern.
- Machine learning models can predict PFAS exposure risks using environmental data.
- Data preprocessing methods can influence model outcomes, but their effects are not well understood.
Purpose of the Study:
- To evaluate how different data preprocessing techniques affect machine-learned Bayesian network models for predicting perfluorooctanoic acid (PFOA) in groundwater.
- To assess the impact of data preprocessing on model accuracy and the identification of high-risk exposure areas.
Main Methods:
- Utilized 19 years of PFOA measurements from Minnesota, USA, linked to potential sources and environmental fate factors.
- Applied nine distinct data preprocessing methods to create training datasets.
- Trained Bayesian network models to predict the probability of PFOA concentrations exceeding the Minnesota health advisory level (35 ppt).
Main Results:
- Varying preprocessing approaches resulted in different model structures and accuracies.
- All models demonstrated robust performance, accurately distinguishing between high-risk and low-risk wells (82.0%–89.0% accuracy).
- A trade-off exists between data quality and model performance; stricter screening reduced sample size but potentially improved data quality.
Conclusions:
- Data preprocessing is a critical step that influences the structure and accuracy of machine learning models for PFAS exposure assessment.
- Despite preprocessing variations, the developed models effectively identified groundwater wells at high risk for PFOA contamination.
- Optimizing data preprocessing strategies is essential for reliable PFAS risk prediction, balancing data quantity with quality.
More Related Videos
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
07:06Investigating Long-Distance Transport of Perfluoroalkyl Acids in Wheat via a Split-Root Exposure Technique
Published on: September 28, 2022
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model Approaches for Pharmacokinetic Data: Physiological Models
Mechanistic Models: Compartment Models in Individual and Population Analysis
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
