Related Experiment Video
Updated: Jul 13, 2026

E-Patient Counseling Trial E-PACO: Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Artificial intelligence performance in generating colorectal surgery board questions
Jonathan Zuo1, Makenna Marty1, Seija Maniskas1
1Division of Colorectal Surgery, Huntington Health - an Affiliate of Cedars-Sinai, Pasadena, CA, USA.
Background:
Large language models (LLM) can pass medical licensing and specialty board exams, but their ability to generate high-quality board-style exam questions is uncertain.
Methods:
Three LLMs each generated 20 colorectal surgery board questions in accordance with American Board of Colon and Rectal Surgery guidelines. Questions from the Colon and Rectal Surgery Educational Program (CARSEP) served as comparators. Board-certified colorectal surgeons, blinded to source, graded each question on clarity, relevance, suitability, distractor quality, and adequacy of rationale, and categorized questions as "Approved for Committee," "Author to Review," or "Not Accepted."
Results:
CARSEP demonstrated the highest "Approved for Committee" rate (65%), compared with ChatGPT-4o (7%), Copilot Pro (10%), and Gemini Advanced (10%). CARSEP significantly outperformed most LLMs across all evaluation domains (p < 0.001), except for question suitability, where most LLM questions received >70% very good to good ratings.
Conclusions:
Although LLMs demonstrate potential, they are currently unable to consistently generate high-quality, colorectal surgery board-style questions.

