Related Experiment Video
Updated: Feb 17, 2026

07:13
Digital Home-Monitoring of Patients after Kidney Transplantation: The MACCS Platform
Published on: April 12, 2021
5.3K
Patient Education in Bariatric Surgery: Can Artificial Intelligence-Based Chatbots Bridge the Knowledge Gap?
Amirreza Izadi1,2, Hesam Mosavari1,2, Ali Hosseininasab1,2
1Department of Surgery, Surgery Research Center, School of Medicine, Rasool-E Akram Hospital, Iran University of Medical Sciences, Tehran, Iran, iums.ac.ir.
Journal of Obesity
|February 16, 2026
Summary
AI chatbots provided accurate answers to bariatric surgery questions, outperforming human experts. While promising for patient education, readability and accuracy require improvement for clinical use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Surgical Education
Background:
- The global obesity epidemic necessitates effective patient education for metabolic and bariatric surgery (MBS).
- Limited resources in MBS centers create knowledge gaps, prompting patients to seek information online.
- Artificial intelligence (AI) chatbots offer potential for reliable medical information dissemination, but accuracy concerns persist.
Purpose of the Study:
- To evaluate the accuracy and comprehensiveness of AI chatbots in answering patient questions about laparoscopic sleeve gastrectomy (LSG).
- To compare the performance of various AI chatbots against a multidisciplinary expert team in providing LSG information.
- To assess the readability and reliability of AI-generated responses for patient education.
Main Methods:
- Seven AI chatbots (ChatGPT 3.5/4, Bard, Bing, Claude, Llama, Perplexity) were tested.
- Forty LSG-related questions were sourced from patient forums and social media.
- Responses from chatbots and a multidisciplinary expert team (surgeons, fellows, GPs) were evaluated for accuracy and comprehensiveness on a 5-point scale.
Main Results:
- AI chatbots achieved higher overall performance scores (2.55) than the expert group (1.92; p < 0.001).
- ChatGPT-4 demonstrated the highest performance among chatbots (2.94), while Llama had the lowest (2.15).
- Chatbot responses generally required an 11th-grade to college reading level, with ChatGPT-4 showing the best reliability.
Conclusions:
- AI chatbots show promise as a scalable tool for bariatric patient education, generating accurate and comprehensive answers.
- Readability, model variability, occasional inaccuracies, and medicolegal aspects require further attention.
- Chatbots should supplement, not replace, clinician counseling, with future research focusing on improving readability and real-world impact.
