Related Experiment Video
Updated: Mar 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the clinical competence of large language models for tobacco use disorder: A multi-domain expert evaluation
Thiago P Fernandes1, Linnea Dahlgren2, Natanael A Santos1,3
1Department of Psychology, Perception, Neuroscience, and Behaviour Lab, Federal University of Paraiba, Joao Pessoa.
Background:
Tobacco use disorder (TUD) remains the leading preventable cause of death globally, yet fewer than one-third of users receive guideline-concordant care due to workforce shortages and training gaps. Emerging artificial intelligence (AI) systems, particularly large language models (LLMs), may help expand access to evidence-based cessation support, but their clinical competence remains insufficiently characterized.
Objective:
To systematically evaluate the clinical accuracy, safety, guideline adherence, and communication quality of five leading LLMs across standardized tobacco-cessation scenarios.
Methods:
We developed 84 clinical vignettes, covering five core competency domains (screening, diagnosis, pharmacotherapy, behavioral counseling, and harm reduction) based on DSM-5-TR, U.S. Public Health Service Guidelines, NICE NG209, ASAM guidelines, SAMHSA protocols, and the WHO MPOWER framework. Five LLMs, GPT-4.5, Claude 3.5 Sonnet, Gemini 2.5 Pro, Llama 3.1-70B, and DeepSeek-V3, were independently evaluated by four addiction-medicine experts across clinical accuracy, guideline adherence, safety, and clinical utility.
Results:
GPT-4.5 and Claude 3.5 Sonnet achieved the highest composite scores (M = 4.25, SD = 0.68; M = 4.18, SD = 0.71), with 74-78% of responses rated ≥4.0 and superior safety performance (88% ≥4.0). Gemini 2.5 Pro showed moderate performance (M = 3.72; 52% ≥4.0). Open-weight models (Llama 3.1-70B, M = 3.48; DeepSeek-V3, M = 3.35) lagged behind overall, with only 34-38% achieving the ≥4.0 benchmark.
Conclusion:
All evaluated AI systems demonstrated competence in tobacco-cessation counseling, but GPT-4.5 and Claude 3.5 Sonnet reached performance levels consistent with supervised clinical use and safety-critical scenarios. Clinician oversight remains essential for all medication-based interventions, and open-weight models warrant further validation before consideration for clinical implementation.
More Related Videos
05:12Chronic Intermittent Ethanol Vapor Exposure Paired with Two-Bottle Choice to Model Alcohol Use Disorder
Published on: June 23, 2023
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Related Concept Videos
Drug Dependence
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Substance Use Disorders Affecting Sleep
Understanding the concepts of physical dependence,...
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Diagnostic and Statistical Manual of Mental Disorders (DSM)
CNS Depressants: Alcohol and Nicotine