Related Experiment Video
Updated: Jun 19, 2026

10:42
A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
6.9K
Can Large Language Models Be a Viable Tool for Consensus Working Groups? Experience of the Ventral Rectopexy Expert
Frank G Lee1, Ellen L Larson1, Jane Vermunt1
1Division of Colon and Rectal Surgery, Mayo Clinic, Rochester, Minnesota.
Diseases of the Colon and Rectum
|January 6, 2026
Summary
OpenEvidence demonstrated superior content appropriateness and citation quality compared to Gemini and ChatGPT in a study on ventral rectopexy literature synthesis. Experts could not distinguish chatbot-generated text from human consensus.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Surgical Literature Synthesis
Background:
- A recent consensus update on ventral rectopexy was published by the Ventral Rectopexy International Expert Panel.
- The study evaluated the capability of large language models (LLMs) to synthesize ventral rectopexy literature before the consensus publication.
Purpose of the Study:
- To compare the performance of different LLMs, including ChatGPT-4o, Gemini 1.5 Pro, and OpenEvidence, in generating responses and citations related to ventral rectopexy.
- To use the expert panel consensus as a benchmark for evaluating LLM performance.
Main Methods:
- LLMs were assessed for content appropriateness, readability, response length, citation fabrication, and citation quality.
- A panel of colorectal surgeons evaluated the LLM-generated text for appropriateness and attempted to distinguish it from the expert consensus.
Main Results:
- OpenEvidence achieved the highest content appropriateness (3.5/5), surpassing Gemini (3.0/5) and ChatGPT (2.8/5).
- ChatGPT exhibited the highest verbosity and readability, but fabricated 53% of its citations, compared to 12% for Gemini and 0% for OpenEvidence.
- Colorectal surgeons could identify chatbot-generated text with 55% accuracy.
Conclusions:
- OpenEvidence excelled in content appropriateness and citation quality over Gemini and ChatGPT.
- LLM-generated text was indistinguishable from expert consensus to specialists.
- LLMs may serve as valuable research tools for future consensus groups with proper disclosure and oversight.
Related Concept Videos
Stereotype Content Model
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence categorization, a person will feel...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

