Related Experiment Video
Updated: Jun 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Characterizing Parental Use of a HIPAA-Compliant Large Language Model Chatbot in Rare Pediatric Diseases: Insights
Anna M Kerr1, Christine Bereitschaft1, Stephanie Milam1
1Washington University School of Medicine, Division of Pediatric Hematology and Oncology, Department of Pediatrics, United States.
Background:
Parents of children with rare and serious illnesses often have unmet information needs. Large language models (LLMs) can help parents seek medical information. However, few studies have observed parents' use of LLMs or how they would use it in conjunction with their patient portal.
Objectives:
We provided parents of children with cancer or vascular anomalies (VAs) access to a secure Health Insurance Portability and Accountability Act (HIPAA)-compliant chatbot. We characterized how parents used the tool while accessing their child's patient portal and evaluated the chatbot responses.
Methods:
Parents participated in think-aloud sessions (n = 48). Parents accessed a HIPAA-compliant GPT 4 Endpoint and entered queries about their child's illness. We examined query length and chatbot response length, accuracy, and readability. We also conducted content analysis on parent queries.
Results:
We analyzed 451 queries and 451 responses. Parents' queries ranged from 1 to 104 words. They entered primarily short well-formed questions or phrases/statements. Some entered single words or incomplete phrases. Content was related to diagnosis/etiology, treatment, symptoms/side effects, laboratory values, imaging results, clinician notes/documentation, and supportive resources, with some differences between VA and cancer contexts. Chatbot responses ranged from 9 to 883 words The mean accuracy rating was 4.9 ± 0.5 and the mean Flesch Reading Ease score was 28.4 ± 15.0 (college-graduate level).
Conclusion:
Parents' queries varied in length, complexity, and content, with some differences indicating unique information needs by disease context. Chatbot responses were accurate yet written at a reading level potentially challenging for some users. Future studies should consider these patterns and characteristics when designing health-related chatbot-based tools.
