Related Experiment Video
Updated: Jun 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
ChatGPT Versus Custom-Trained Chatbot for Urogynecology Surgery Counseling
Leanne Brechtel1, Colin Johnson1, Sarah Rabice1
1Division of Urogynecology and Reconstructive Pelvic Surgery, Department of Obstetrics and Gynecology, University of Iowa Hospitals and Clinics, Iowa City, IA.
Importance:
Patient education materials are important components of shared decision making in urogynecologic surgery. Traditional materials are often difficult to update, lack personalization, and are not easily adaptable. Chatbots offer a new approach to generating comprehensive and accessible content, but their utility in a clinic setting remains unclear.
Objective:
The objective of this study was to compare the performance of a general-purpose chatbot, ChatGPT, and a domain-specific chatbot developed by the Foundation for Female Health Awareness (FFHA) (FFHA Assistant) in generating surgical counseling information, using standardized materials from the International Urogynecological Association (IUGA) as reference.
Study Design:
Seven IUGA handouts representing common urogynecologic surgical procedures were selected. Identical prompts were submitted to ChatGPT-4.0 and the FFHA Assistant. Responses were reviewed by 7 blinded urogynecology experts using 5-point Likert scales to assess accuracy, completeness, and understandability. Readability was evaluated using the Flesch-Kincaid Grade Level and Flesch Reading Ease Score.
Results:
ChatGPT-4.0 outperformed in completeness as compared with the IUGA leaflets (median 4 [3-5] vs 3 [3-4], P<0.01), whereas the FFHA Assistant scored higher in accuracy (median 3 [3-3] vs 3 [2-3], P<0.01) and understandability (3 [3-4] vs 3 [3-3], P<0.01). Both large language models generated longer responses than the IUGA leaflets. The FFHA Assistant responses had better readability scores, aligning more closely with health literacy recommendations.
Conclusions:
Both chatbots generated counseling content comparable to or superior to existing materials. The domain-specific FFHA Assistant responses were better aligned with health literacy recommendations. Further research is needed to better understand the reproducibility of responses and the clinical utility of chatbots in patient education.
Related Concept Videos
Urologic Endoscopic Procedure: Cystoscopic Examination
Urinary Tract Calculi VI: Surgical Management
