Related Experiment Video
Updated: Sep 14, 2025

Murine Drinking Models in the Development of Pharmacotherapies for Alcoholism: Drinking in the Dark and Two-bottle Choice
Published on: January 7, 2019
Stigmatizing Language in Large Language Models for Alcohol and Substance Use Disorders: A Multimodel Evaluation and
Yichen Wang1, Kelly Hsu, Christopher Brokus
1Division of Hospital Medicine, Department of Medicine, Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA (YW); Gastroenterology Unit, Department of Medicine, Massachusetts General Hospital, Boston, MA (YW, NU, WZ); Tufts University School of Medicine, Boston, MA (KH); Harvard Medical School, Boston, MA (CB, NU, WZ); Division of Gastroenterology and Hepatology, Department of Medicine, Mayo Clinic, Jacksonville, FL (YH); Program for Substance Use and Addiction Services, Massachusetts General Hospital, Boston, MA (SW); Department of Biomedical Data Science, Stanford University, Stanford, CA (JZ); Department of Electrical Engineering, Stanford University, Stanford, CA (JZ); Department of Computer Science, Stanford University, Stanford, CA (JZ).
Objectives:
Large language models (LLMs) are increasingly used in health care communication but can inadvertently perpetuate stigmatizing language toward individuals with alcohol and substance use disorders. Despite growing interest in LLM performance, a focused evaluation of their propensity for SL and strategies to mitigate it remains lacking.
Methods:
We generated 60 clinically relevant questions ["prompts"; 20 each for alcohol use disorder (AUD), alcohol-associated liver disease (ALD), and substance use disorder (SUD)] and tested 14 LLMs. Two physicians independently assessed all responses for stigmatizing language using guidelines from the National Institute on Drug Abuse and the National Institute on Alcohol Abuse and Alcoholism; discrepancies were resolved by a third physician. We employed iterative prompt engineering (PE)-a process of strategically crafting input instructions to guide model outputs towards nonstigmatizing language-to reduce stigmatizing language by incorporating a list of specific terms to avoid and identifying model-specific pitfalls. We compared the prevalence of SL in responses to native prompts (baseline, unengineered) versus engineered prompts, adjusting for word count in multivariate analyses.
Results:
Of 840 responses generated from native prompts, 297 (35.4%) contained stigmatizing language, totaling 592 terms. With prompt engineering, only 53 (6.3%) of 840 responses contained stigmatizing language, comprising 104 terms. Prompts on topic of ALD yielded higher odds of stigmatizing language than those addressing AUD (adjusted odds ratio, 2.11; 95% CI, 1.47-3.02; P < 0.001), whereas prompts on substance use disorder (SUD) did not differ significantly from AUD (adjusted odds ratio, 1.17; 95% CI, 0.81-1.69; P = 0.40). Prompt engineering reduced the likelihood of stigmatizing language by 88% in univariate analysis ( P < 0.001), and this effect persisted after adjusting for word count (adjusted odds ratio, 0.15; 95% CI, 0.11-0.20; P < 0.001).
Conclusions:
LLMs frequently generated stigmatizing language when discussing alcohol-related and substance use-related conditions, potentially undermining patient-centered care. However, targeted prompt engineering substantially reduced stigmatizing language occurrences across diverse models. These findings emphasize the need for ongoing model refinement and structured prompting strategies to ensure stigma-free language in health care communication.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Stereotype Content Model
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Self-Presentation: Self-Monitoring and Self-Handicapping

