Related Experiment Video
Updated: Aug 13, 2026

07:14
Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Finite State Machine-Guided Retrieval-Augmented Generation Improves Expert-Rated Acceptability of a Peripherally
Mangyeong Lee1,2, Seung-Beom Cho3, Jae-Wook Yu3
1Center for Clinical Epidemiology, Samsung Medical Center, Seoul, Republic of Korea.
Journal of Medical Internet Research
|August 11, 2026
Summary
A finite state machine-guided retrieval-augmented generation (RAG) model improved chatbot acceptability for peripherally inserted central catheter (PICC) management. Expert nurses preferred this advanced model over simpler versions for clinical decision-making support.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Decision Support Systems
- Natural Language Processing for Medical Applications
Background:
- Patients undergoing chemotherapy often require peripherally inserted central catheters (PICCs) for treatment administration.
- Self-management of PICCs is crucial for ambulatory cancer patients to ensure treatment continuity.
- Large language models (LLMs) offer potential for patient support but face challenges with accuracy and protocol adherence.
Purpose of the Study:
- To evaluate if a finite state machine (FSM)-guided retrieval-augmented generation (RAG) architecture enhances clinical acceptability of chatbots for PICC management.
- To compare the FSM-guided RAG model against simpler chatbot architectures (fine-tuned model alone, fine-tuned model with RAG).
Main Methods:
- A blinded, comparative evaluation of three chatbot architectures based on a fine-tuned GPT-4o-mini model.
- Three expert PICC nurses evaluated chatbot responses to 43 queries using forced-choice preference and Likert scale ratings.
- Statistical analysis included Gwet AC1 for interrater agreement and Cochran Q/McNemar tests for preference comparison.
Main Results:
- The FSM-guided RAG model (model 3) was preferred by expert nurses in 79.1% of scenarios (34/43).
- Model 3 showed significantly higher preference rates compared to the fine-tuned model alone (model 1) and the FT+RAG model (model 2).
- While generally rated higher, the FSM-guided model received slightly lower efficiency ratings, attributed to its conversational structure.
Conclusions:
- The FSM-guided fine-tuned RAG model demonstrated superior expert-rated acceptability for PICC management support.
- FSM-based dialogue control effectively aligns chatbot outputs with expert clinical consultation patterns.
- Further patient-facing usability testing and multicenter validation are necessary before clinical deployment.