Related Experiment Video
Updated: Sep 13, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Comparative analysis of AI enhancements to rules-based clinical triage chatbot: Convergent mixed-methods study
Aleksa Stankić1, Dušan Vujošević1, Nemanja Radosavljević1
1Union University School of Computing, Department of Computer Engineering, Belgrade, Serbia.
Abstract:
BackgroundChatbots are increasingly used in healthcare for symptom checking and care navigation. However, concerns about accuracy limit large language model (LLM) use in clinical settings.ObjectiveThis convergent mixed-methods study evaluated three enhancements to an existing rules-based triage chatbot. The purpose is to inform the development of tools improving patient access and system navigation.MethodsThree artificial intelligence tools (Dialogflow, ChatGPT, and Gemini) were integrated into a rules-based workflow from December 21st 2024 to July 7th 2025. Thirteen testers completed 156 vignette-based interactions and 52 Likert-scale evaluations from July 17th 2025 to August 5th 2025. Quantitative outcomes included accuracy, conversational turns, time to completion, and tester ratings. Qualitative feedback was thematically analyzed.ResultsThe rules-based chatbot was most accurate (71.8%, 95% CI 56.2-83.5%), outperforming Dialogflow (30.8%, p = 0.002) and Gemini (41.0%, p = 0.031). Dialogflow, ChatGPT, and Gemini were statistically indistinguishable. Qualitative analysis revealed user preference for ChatGPT and Gemini despite equal or worse performance than the rules-based chatbot.ConclusionRules-based triage remained most reliable, and the divergence between user preference and measured accuracy underscores that safe clinical deployment of LLM-based chatbots requires governance addressing accuracy, interpretability, and masking of poorer performance by conversational fluency.