Related Experiment Video
Updated: Sep 17, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Identifying Key Predictors of Smoking Cessation Success: Text-Based Feature Selection Using a Large Language Model
Thuy T T Le1, Jiongxuan Yang2, Zimo Zhao3
1University of Michigan School of Public Health, Department of Health Management and Policy, Ann Arbor, MI, USA.
Background:
The most effective way to reduce mortality and morbidity among current smokers is to quit smoking. Although about half of smokers attempted to quit, only one-tenth succeeded in 2022.
Objective:
To identify key predictors of smoking cessation success to inform cessation interventions and increase quitting rates.
Methods:
We analyzed data from waves 5 and 6 of the Population Assessment of Tobacco and Health (PATH) study (December 2018 to November 2021). Using OpenAI's GPT-4.1, we identified the top 45 variables from wave 5 that are highly predictive of 12-month smoking abstinence in wave 6, based on descriptions of survey variables. We then validated the predictive power of the GPT-4.1-selected variables by comparing the performance of eXtreme Gradient Boosting (XGBoost) trained on different sets of variables. Finally, we derived insights into the top 10 variables, ranked according to their SHapley Additive exPlanations values.
Results:
The performance of XGBoost trained with all possible wave 5 variables and the 45 selected variables was almost identical (AUC:0.749 vs AUC:0.752). The top 10 variables included past 30-day smoking frequency, minutes from waking up to smoking first cigarette, important people's views on tobacco use, prevalence of tobacco use among close associates, daily electronic nicotine product use, emotional dependence, and health harm concerns.
Conclusion:
This study demonstrates the ability of OpenAI's GPT-4.1 to identify the top 45 PATH wave 5 variables associated with 12-month smoking abstinence using only their descriptions. This approach could help researchers design more effective survey questionnaires and improve efficiency of data collection.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
Longitudinal Research
Quantifying and Rejecting Outliers: The Grubbs Test
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Expected Frequencies in Goodness-of-Fit Tests

