Related Experiment Video
Updated: Apr 29, 2026

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
Understanding online sentiment toward waterpipe tobacco smoking by applying deep-learning language models to Twitter
Zidian Xie1, Puhua Ye2, Mengwei Wu2
1Clinical and Translational Science Institute, University of Rochester Medical Center, Rochester, NY 14642, United States.
Introduction:
Waterpipe tobacco smoking remains a popular social activity among young adults in the United States. This study aims to understand sentiment toward waterpipe on social media in the United States by analyzing Twitter/X data.
Methods:
Using keywords ("hookah," "waterpipe," "shisha"), we collected US tweets posted between March 2021 and March 2023. Commercial content (eg, containing "sale," "discount," "$") was filtered out, yielding 299 544 non-commercial tweets. A random sample of 2300 tweets was manually coded for sentiment (positive/negative/neutral) and for whether the author might use waterpipe, which were used to fine-tune a deep-learning model (Llama-2). We applied BERTopic modeling to identify main themes.
Results:
Overall, tweets with a positive sentiment (57.0%, 170 597/299 544) were higher than those with a negative sentiment (16.7%, 50 196/299 544). Among those Twitter users who might use waterpipe, 82.0% of their tweets showed a positive sentiment toward waterpipe while only 7.5% of the tweets had a negative sentiment. In contrast, among those who might not use waterpipe, 29.6% of their posts showed a positive sentiment and 27.0% of the posts with a negative sentiment. Main themes identified from positive tweets included excitement about waterpipe lounges, enjoyment of specific flavors, and social desires for waterpipe use. Negative tweets focused on health concerns of waterpipe, discomfort from waterpipe smoking, criticism of promotional content, and negative social sentiments.
Conclusion:
Our results provide preliminary insights into how waterpipe smoking was perceived and discussed among Twitter users in the United States, which could help with future targeted social media-based public health intervention campaigns.
Implications:
By applying fine-tuned Llama-2 language models and BERTopic modeling, this study showed how the public perceived waterpipe on Twitter, which varied depending on the user status (whether they used waterpipe or not), time of day or day of week, and geolocation. These findings offer some insights into the waterpipe discourse on Twitter. More importantly, our understanding of different sentiments toward waterpipe will provide useful guidance for designing effective health communication to reduce or prevent waterpipe use.
