Related Experiment Video
Updated: Sep 2, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Evolution of Generative Artificial Intelligence in Clinical Practice: Comparative Performance of OpenEvidence 2.0 and
Fedan Avrumova1, Michael D D Cesar2, Giuseppe Loggia3
1Department of Spine Surgery, Hospital for Special Surgery, New York, NY, USA.
Background Context:
Generative artificial intelligence (AI) is increasingly used in spine care; however, concerns remain regarding citation hallucinations and reliability. ChatGPT may generate inaccurate or fabricated references, whereas OpenEvidence (OE) prioritizes verified, peer-reviewed literature. This is the first study comparing OE and ChatGPT using cervical spine clinical guideline (CSCG) queries.
Purpose:
To compare guideline alignment, citation validity, sourcing, and prompt-engineering effects between OE and ChatGPT using CSCGs.
Study Design/Setting:
Cross-sectional comparative analysis.
Patient Sample:
No patient population was included.
Outcome Measures:
Primary outcomes were guideline alignment score and citation validity (fully correct, partially hallucinated, or fully hallucinated). Secondary outcomes included source type, publication year, proportion published after CSCG release, and prompt-engineering effects.
Methods:
A total of 110 evidence-based clinical questions derived from 10 CSCGs authored by 6 academic societies were submitted to OE (v2.0) and ChatGPT-4o from June 1 to July15, 2025. A subset of prompts was repeated to evaluate prompt-engineering effects.
Results:
OE generated 999 citations with 100% accuracy, whereas only 184/393 (46.8%) ChatGPT citations met accuracy criteria (p<0.001). OE demonstrated higher guideline alignment than ChatGPT (4.6 ± 0.8 vs 4.1 ± 0.7; p = 0.03), with almost perfect interrater agreement (weighted Cohen's κ = 0.88; 95% CI, 0.80-0.97). . ChatGPT produced 88 partially hallucinated citations (22.4%), most commonly due to incorrect hyperlinks, author names, or publication years. OE cited more peer-reviewed literature (79.8% vs 60.1%; p<0.001) and more recent studies (2018±5.6 vs 2012±7.4; p<0.001). Prompt-engineering analysis showed OE maintained higher citation validity and fewer hallucinations, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
Conclusion:
OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.
