Related Experiment Video
Updated: Sep 2, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Evolution of generative artificial intelligence in clinical practice: comparative performance of OpenEvidence 2.0 and
Fedan Avrumova1, Michael D D Cesar2, Giuseppe Loggia3
1Department of Spine Surgery, Hospital for Special Surgery, New York, NY, USA.
Background Context:
Generative artificial intelligence (AI) is increasingly used in spine care; however, concerns remain regarding citation hallucinations and reliability. ChatGPT may generate inaccurate or fabricated references, whereas OpenEvidence (OE) prioritizes verified, peer-reviewed literature. This is the first study comparing OE and ChatGPT using cervical spine clinical guideline (CSCG) queries.
Purpose:
To compare guideline alignment, citation validity, sourcing, and prompt-engineering effects between OE and ChatGPT using CSCGs.
Study Design/Setting:
Cross-sectional comparative analysis.
Patient Sample:
No patient population was included.
Outcome Measures:
Primary outcomes were guideline alignment score and citation validity (fully correct, partially hallucinated, or fully hallucinated). Secondary outcomes included source type, publication year, proportion published after CSCG release, and prompt-engineering effects.
Methods:
A total of 110 evidence-based clinical questions derived from 10 CSCGs authored by 6 academic societies were submitted to OE (v2.0) and ChatGPT-4o from June 1 to July15, 2025. A subset of prompts was repeated to evaluate prompt-engineering effects.
Results:
OE generated 999 citations with 100% accuracy, whereas only 184/393 (46.8%) ChatGPT citations met accuracy criteria (p<.001). OE demonstrated higher guideline alignment than ChatGPT (4.6±0.8 vs 4.1±0.7; p=.03), with almost perfect interrater agreement (weighted Cohen's κ = 0.88; 95% CI, 0.80-0.97). ChatGPT produced 88 partially hallucinated citations (22.4%), most commonly due to incorrect hyperlinks, author names, or publication years. OE cited more peer-reviewed literature (79.8% vs 60.1%; p<.001) and more recent studies (2018±5.6 vs 2012±7.4; p<.001). Prompt-engineering analysis showed OE maintained higher citation validity and fewer hallucinations, although accuracy declined when outputs were reformatted through ChatGPT, suggesting cross-model contamination.
Conclusion:
OE demonstrated superior CSCG concordance and substantially lower susceptibility to hallucinations than ChatGPT. Nonetheless, physician oversight remains essential for safe AI integration into clinical practice.
