Related Experiment Video
Updated: Aug 6, 2026

E-Patient Counseling Trial (E-PACO): Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Correctness, Harmfulness, and Diversity of Large Language Models for Colonoscopy Preparation Assistance: Comparative
Tomiris Kaumenova1, Subhankar Chakraborty2, Eric Fosler-Lussier1,3
1Department of Linguistics, The Ohio State University, 1712 Neil Ave, Columbus, OH, 43210, United States.
Large language models (LLMs) show promise for improving colonoscopy preparation by answering patient questions. However, current models produce harmful errors and are not yet reliable for direct patient use.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing
- Medical Informatics
Background:
- Colorectal cancer is a major health concern, with colonoscopy being crucial for early detection.
- Inadequate bowel preparation frequently leads to postponed colonoscopies due to patient difficulties with instructions.
- Existing digital tools offer limited support for patient-specific preparation questions.
Purpose of the Study:
- To evaluate the correctness, safety, and diversity of AI-generated dialogues for colonoscopy preparation.
- To compare the performance of leading large language models (LLMs) in simulating patient-AI Coach interactions.
- To assess the effectiveness of a safety filter in mitigating errors during AI-driven patient support.
Main Methods:
- Five leading LLMs generated 250 simulated patient-AI Coach dialogues each.
- Dialogues covered diet, medications, and preparation instructions, using a multiprompt approach for question diversity.
- Human experts and an LLM-as-a-judge system evaluated dialogue correctness, error types, and potential harmfulness.
Main Results:
- Leading LLMs showed potential but did not achieve adequate performance for colonoscopy preparation.
- Closed-weight models (GPT-5.1, GPT-4.1, o3) outperformed open-weight models (Llama, Mistral) in response correctness.
- All models generated harmful errors; a safety filter reduced but did not eliminate these issues.
Conclusions:
- LLMs offer potential for colonoscopy preparation support but require further development for safe deployment.
- Addressing persistent harmful errors and improving safety mechanisms are critical for patient-facing applications.
- Validation with real patient queries is essential before widespread use of LLM-based preparation tools.
Related Concept Videos
Imaging Studies III: Gastrointestinal Motility Studies and Virtual Colonoscopy
Radionuclide Testing
Radionuclide testing is a sophisticated medical technique for assessing gastrointestinal motility. It focuses on gastric emptying and colonic transit time. Radioactive markers track the movement of food through the digestive system, providing insights into gastrointestinal disorders.
In gastric emptying studies, a meal's liquid and solid...
Endoscopic Procedures II: Colonoscopy
