Related Experiment Video
Updated: Aug 13, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Finite State Machine-Guided Retrieval-Augmented Generation Improves Expert-Rated Acceptability of a Peripherally
Mangyeong Lee1,2, Seung-Beom Cho3, Jae-Wook Yu3
1Center for Clinical Epidemiology, Samsung Medical Center, Seoul, Republic of Korea.
Background:
Patients with cancer undergoing long-term or vesicant chemotherapy frequently require peripherally inserted central catheters (PICCs). Due to the nature of ambulatory treatment administration, self-PICC management is essential for the continuation and completion of the planned treatment. Large language models offer potential for continuous patient support, but hallucinations and insufficient adherence to clinical protocols remain concerns. Fine-tuning (FT) and retrieval-augmented generation (RAG) improve factual grounding but cannot enforce the structured decision logic of expert-led PICC consultations, leaving the value of dialogue-control mechanisms unclear.
Objective:
This study is an expert-panel evaluation designed to exploratorily verify whether a finite state machine (FSM)-guided RAG architecture yields incremental gains in clinical acceptability when added on top of FT and RAG compared with simpler architectures (fine-tuned model alone; fine-tuned model with RAG).
Methods:
We conducted a blinded ablation comparison of three chatbot architectures sharing an identical fine-tuned GPT-4o-mini base: model 1 (fine-tuning only; FT), model 2 (FT + RAG), and model 3 (FT + RAG + FSM). Three nurses specializing in PICC management (≥10 y experience) independently evaluated responses to 43 PICC-related queries. The two-stage evaluation comprised query-level forced-choice preference and global Likert ratings across 8 predefined dimensions. Interrater agreement was quantified with Gwet AC1. Preference rates were compared with Cochran Q and pairwise McNemar tests with Bonferroni correction. Global Likert ratings were summarized descriptively across 8 predefined dimensions.
Results:
The FSM-guided model (model 3) was preferred in 34 of 43 scenarios (79.1%). Interrater agreement was moderate (Gwet AC1=0.58, 95% CI 0.42-0.75; P<.001). Differences in model preference were significant (Cochran Q=48.2; P<.001). Pairwise McNemar tests showed model 3 was preferred significantly more often than model 1 and model 2 (both P<.001), while model 1 versus model 2 did not differ after Bonferroni correction. For the global ratings, descriptive statistics indicated that model 3 received higher ratings than models 1 and 2 across most criteria, although it tended to receive lower ratings for efficiency, possibly reflecting its multiturn structure. In qualitative debriefing, the nurses noted occasional unnecessary conversational turns under FSM guidance.
Conclusions:
In this expert-panel content-validation study, the FSM-guided fine-tuned RAG model received higher expert-rated acceptability than the simpler architectures, with FSM-based dialogue control more closely aligning chatbot outputs with expert-led PICC consultation patterns. Patient-facing usability testing and multicenter validation are required before claims regarding patient safety or clinical deployment can be made.