Related Experiment Video
Updated: Jun 16, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Dataset development and performance analysis of a specialized conversational artificial intelligence model for
Ebrahim Mellatdoust Pordel1, David Akopian1, Patricia Chalela2
1Department of Electrical & Computer Engineering, The University of Texas at San Antonio, One UTSA Circle, San Antonio, TX 78249, United States.
None:
Conversational artificial intelligence (AI) is increasingly being used in healthcare to encourage behavioral modification, information dissemination, and patient engagement. However, its effectiveness depends on high-quality, domain-specific training data. General-purpose datasets are often inadequate in terms of clinical specificity and dialog coherence, particularly in sensitive domains like smoking cessation. This study introduces a curated question-and-answer dataset comprising 3000 Q&A pairs from reputable health resources, including expert-authored guidelines and cessation programs. Each entry is annotated with metadata reflecting informational intent, action-oriented content, and user engagement signals, supporting both single-turn and multi-turn conversational modeling. A compact generative language model, Chat Generative Pre-trained Transformer 4o-mini (ChatGPT-4o-mini), was fine-tuned using this dataset and compared against the pretrained base model using BLEU, ROUGE-L, and BERTScore metrics, assessing syntactic accuracy, lexical coverage, and semantic alignment. The fine-tuned model demonstrated improved response clarity, conversational coherence, and adherence to evidence-based content. This study emphasizes methodological benchmarking and model performance rather than behavioral or clinical outcomes, isolating the model-optimization layer to assess how curated, domain-specific training data influence response quality under controlled conditions. By contributing a reproducible pipeline for dataset construction and evaluation, this work supports safe and effective conversational agents in healthcare contexts and promotes responsible AI practices in digital health applications. Statement of significance The success of large language model (LLM)-based conversational agents depends on access to high-quality, domain-specific training data. However, such data is typically absent from models trained on general web corpora. Smoking cessation, as a behavior-change domain with substantial clinical burden and health disparities, requires emotionally attuned and evidence-based messaging to be effective. Rather than introducing a new smoking cessation intervention, this study contributes a reproducible methodological framework for evaluating how curated, evidence-based datasets influence LLM behavior in health education contexts. By isolating the dataset-to-model performance relationship, this work provides a technical foundation that complements prior user-centered intervention studies and supports responsible development of conversational AI systems prior to behavioral deployment. This study contributes a rigorously curated dataset drawn from evidence-based sources, annotated with metadata for motivational framing, tone, and dialogue intent. In parallel, we present a methodology for fine-tuning LLMs and evaluating performance in both single-turn and multi-turn conversational settings.