Related Experiment Video
Updated: Apr 30, 2026

15:49
Flexible Colonoscopy in Mice to Evaluate the Severity of Colitis and Colorectal Tumors Using a Validated Endoscopic Scoring System
Published on: October 16, 2013
31.7K
Precision of Chatbot Generative Pretrained Transformer Version 4-Generated References for Colon and Rectal Surgical
Aaron L Albuck1, Chad M Becnel2, Daniel J Sirna1
1School of Medicine, Tulane University, New Orleans, Louisiana.
The Journal of Surgical Research
|August 9, 2024
Summary
ChatGPT-4 shows potential for generating scientific references but lacks accuracy in details like DOIs and author lists for colon and rectal surgery research.
Area of Science:
- Colorectal Surgery
- Artificial Intelligence in Medicine
- Bibliometrics
Background:
- The increasing use of AI tools like ChatGPT-4 in academic research necessitates evaluating their output quality.
- Assessing the accuracy of AI-generated citations is crucial for maintaining scientific integrity.
Purpose of the Study:
- To evaluate the precision of references generated by ChatGPT-4 for colon and rectal surgery literature.
- To identify specific areas of inaccuracy in AI-generated citations.
Main Methods:
- Ten key terms in colon and rectal surgery were used to prompt ChatGPT-4 for citations.
- Generated references were cross-referenced with Scopus, Google, and PubMed databases by two evaluators.
- Accuracy was assessed based on title, authors, journal, publication year, and Digital Object Identifier (DOI).
Main Results:
- 41% of 100 generated references were fully accurate, but none included a DOI.
- 67% of references were partially accurate (identifiable by title and journal).
- No keyword yielded 100% accuracy across all citation elements; author lists were consistently inaccurate.
Conclusions:
- ChatGPT-4 demonstrates potential for academic literature but requires significant improvement for reliable use in colon and rectal surgery research.
- Inconsistencies in accuracy, missing DOIs, and author attribution errors limit its current applicability.
- Further development is needed before ChatGPT-4 can be fully trusted for generating precise scientific references.

