Related Experiment Video
Updated: Jan 14, 2026

Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Machine Learning for Predicting Environmental Mobility Based on Retention Behavior
Tobias Hulleman1,2, Saer Samanipour1,3,4, Paul R Haddad5
1Queensland Alliance for Environmental Health Sciences (QAEHS), 20 Cornwall Street, Woolloongabba, Brisbane, QLD 4102, Australia.
Abstract:
Very persistent and very mobile (vPvM) substances threaten the environment and human health. These chemicals can persist in aquatic systems and move rapidly due to their affinity for water over soil or other adsorbents. Chemical mobility is usually classified using the organic carbon-water partition coefficient (Koc), but experimental log Koc data are unavailable for most substances. With thousands of new chemicals entering the market annually, there is a growing need for advanced cheminformatics tools to prioritize substances of concern. Because reversed-phase liquid chromatography (RPLC) data are more widely available, they were used here as a proxy for environmental mobility. The organic modifier fraction at elution was applied to assign mobility labels to 146,902 chemicals from an RPLC data set. For each chemical, 881 PubChem fingerprints were computed to capture structural information. A random forest classifier was then trained to predict mobility from retention behavior and fingerprints. The model achieved F1 scores of 0.87, 0.81, and 0.96 for very mobile, mobile, and nonmobile classes, respectively, in the test set. Applied to all REACH-registered chemicals (n = 64,492), the model classified 20% as very mobile, 26% as mobile, and 53% as nonmobile, providing a scalable tool for early identification of vPvM substances.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Steps in Outbreak Investigation
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Migration
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

