Related Experiment Video
Updated: Jan 17, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Can OMFS experts distinguish AI from human manuscripts? A double-blind evaluation using ChatGPT-4
1Department of Oral and Maxillofacial Surgery, Geetanjali Dental and Research Institute, Geetanjali University, Eklingpura, 313001, Udaipur, Rajasthan, India; Department of Dental Research Cell, Dr. D. Y. Patil Dental College and Hospital, Dr. D. Y. Patil Vidyapeeth, Pimpri, 411018, Pune, Maharashtra, India; Department of Oral and Maxillofacial Surgery, Narsinhbhai Patel Dental College and Hospital, Sankalchand Patel University, Visnagar, 384315, Gujarat, India.
Experienced surgeons found AI-generated manuscripts readable but lacking scientific depth and citation accuracy. While AI text is often indistinguishable from human writing, expert oversight is crucial for academic integrity.
Area of Science:
- Oral and Maxillofacial Surgery
- Artificial Intelligence in Academia
- Scientific Publishing
Background:
- Generative AI tools like ChatGPT-4 are increasingly used in academic writing.
- Concerns exist regarding the credibility, scientific depth, and detectability of AI-generated content.
- The ability of experts to discern AI-authored scientific manuscripts requires evaluation.
Purpose of the Study:
- To assess if experienced oral and maxillofacial surgeons (OMFS) can differentiate between AI- and human-authored manuscripts.
- To compare AI- and human-authored manuscripts on coherence, scientific rigor, citation accuracy, and overall quality.
- To evaluate the detectability of AI-generated scientific content.
Main Methods:
- Three core OMFS topics were selected for manuscript generation.
- Manuscripts were independently written by ChatGPT-4 and senior OMFS clinicians.
- Twenty board-certified OMFS reviewers evaluated manuscripts on readability, scientific depth, reference accuracy, writing quality, and rigor, also attempting authorship identification.
Main Results:
- Human-authored manuscripts showed superior scientific depth, reference accuracy, and writing quality compared to AI-generated ones.
- Readability and coherence scores were comparable between human and AI texts.
- Reviewers could not reliably distinguish AI- from human-authored manuscripts (54% accuracy), indicating AI's surface fluency.
Conclusions:
- ChatGPT-4 can produce structurally sound and readable OMFS manuscripts.
- AI-generated content exhibits deficiencies in scientific reasoning and citation accuracy, necessitating expert review.
- Transparent disclosure and editorial safeguards are essential for maintaining scientific integrity as AI integrates into academic workflows.

