Predicting Groundwater PFOA Exposure Risks with Bayesian Networks: Empirical Impact of Data Preprocessing on Model
Runwei Li1,2, Jacqueline MacDonald Gibson2
1Department of Civil Engineering, New Mexico State University, 3035 S Espina St, Las Cruces, New Mexico 88003, United States.
Abstract:
The plethora of data on PFASs in environmental samples collected in response to growing concern about these chemicals could enable the training of machine-learning models for predicting exposure risks. However, differences in sampling and analysis methods across data sets must be reconciled through data preprocessing, and little information is available about how such manipulations affect the resulting models. This study evaluates how data preprocessing influences machine-learned Bayesian network models of PFOA in groundwater. We link 19 years of PFOA measurements from Minnesota, USA, to publicly available information about potential PFOA sources and factors that may influence their environmental fate. Nine different preprocessing methods were tested, and the resulting data sets were used to train models to predict the probability of PFOA ≥ 35 ppt, the 2017 Minnesota health advisory level. Different preprocessing approaches produced varying model structures with significantly different accuracies. Nonetheless, models showed similar relationships between predictor variables and PFOA exposure risks, and all models were relatively accurate, distinguishing wells at high risk from those at low risk for 82.0% to 89.0% of test data samples. There was a trade-off between data quality and model performance since a stricter data screening strategy decreased the sample size for model training.
More Related Videos
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
07:06Investigating Long-Distance Transport of Perfluoroalkyl Acids in Wheat via a Split-Root Exposure Technique
Published on: September 28, 2022
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model Approaches for Pharmacokinetic Data: Physiological Models
Mechanistic Models: Compartment Models in Individual and Population Analysis
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
