Related Experiment Video
Updated: Sep 15, 2025

The Motivation for Alcohol Reward: Predictors of Progressive-Ratio Intravenous Alcohol Self-Administration in Humans
Published on: April 28, 2022
Multivariable machine learning prediction of risky alcohol use in contemporary youth
Lucinda Grummitt1, Rachel Visontay1, Philip Clare2,3,4
1The Matilda Centre for Research in Mental Health and Substance Use, The University of Sydney, Sydney, Australia.
Background And Aims:
Risky alcohol use in young adulthood is a significant public health concern. Understanding the predictors of risky drinking during this period is essential for prevention. This study aimed to measure the predictive accuracy of ensemble machine learning and identify the most important predictors of risky alcohol use in early adulthood.
Design And Setting:
Secondary analysis of the Longitudinal Study of Australian Children, an Australian national longitudinal cohort study.
Participants:
A total of 4983 children, aged 4-5 years in 2004 (Wave 1), followed up for eight waves (to age 18/19 in 2018).
Measurements:
Risky alcohol use was measured at age 18 and defined as more than 10 standard drinks per week, as per Australian National guidelines. Predictors from multiple domains-sociodemographic, adolescent substance use, adolescent mental health and behaviours, parental mental health and substance use, school factors, peer influences, parenting practices and parental stress-were included, measured from Wave 1 to 7. The SuperLearner package in R was used to test a series of models [regularised regression (LASSO, ridge and elastic net), random forest and kernel support vector machine (SVM)] using nested 10-fold cross-validation to identify the overall predictive ability of the model (measured by area under the curve; AUC) and the most important predictors of risky alcohol use across childhood and adolescence. Predictor importance was derived by normalising algorithm-specific scores per fold, weighting them by SuperLearner coefficients and aggregating across folds to rank predictors by mean weighted importance on a scale of 0 to 1 (higher scores indicating greater importance).
Findings:
The ensemble model showed good prediction on the test set, with an AUC of 0.792, a slight improvement over any single algorithm (AUC = 0.783 for the best performing individual algorithm). The most important predictors were weekly drinking at the previous wave (mean weighted importance 0.999), lifetime cannabis use (0.446), lifetime parent financial stress (0.420), identifying as female (0.365), identifying as male (0.344; compared with a reference category of gender diverse), lifetime attention deficit hyperactivity disorder (0.248), pre-natal alcohol exposure (0.248), housing insecurity (0.243), religious involvement (0.238) and parent alcohol use problems (0.215).
Conclusions:
An ensemble learning approach appears to have good predictive ability of risky alcohol use among a contemporary cohort of young Australians. It underscores the complex interplay of individual, familial and social factors occurring across childhood and adolescence that influences risky alcohol use in early adulthood.
More Related Videos
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Predicting Reaction Outcomes
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Hypothesis Test for Test of Independence
H0: The two variables (factors)...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Drug Dependence

