Related Experiment Video
Updated: Jun 8, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
The double-edged sword of generative AI: surpassing an expert or a deceptive "false friend"?
Franziska C S Altorfer1, Michael J Kelly2, Fedan Avrumova2
1Department of Spine Surgery, Hospital for Special Surgery, New York, NY, USA; Universtity Spine Center Zurich, Balgrist University Hospital, University of Zurich, 8006 Zurich, Switzerland.
Background Context:
Generative artificial intelligence (AI), ChatGPT being the most popular example, has been extensively assessed for its capability to respond to medical questions, such as queries in spine treatment approaches or technological advances. However, it often lacks scientific foundation or fabricates inauthentic references, also known as AI hallucinations.
Purpose:
To develop an understanding of the scientific basis of generative AI tools by studying the authenticity of references and reliability in comparison to the alignment of responses of evidence-based guidelines.
Study Design:
Comparative study.
Methods:
Thirty-three previously published North American Spine Society (NASS) guideline questions were posed as prompts to 2 freely available generative AI tools (Tools I and II). The responses were scored for correctness compared with the published NASS guideline responses using a 5-point "alignment score." Furthermore, all cited references were evaluated for authenticity, source type, year of publication, and inclusion in the scientific guidelines.
Results:
Both tools' responses to guideline questions achieved an overall score of 3.5±1.1, which is considered acceptable to be equivalent to the guideline. Both tools generated 254 references to support their responses, of which 76.0% (n=193) were authentic and 24.0% (n=61) were fabricated. From these, authentic references were: peer-reviewed scientific research papers (147, 76.2%), guidelines (16, 8.3%), educational websites (9, 4.7%), books (9, 4.7%), a government website (1, 0.5%), insurance websites (6, 3.1%) and newspaper websites (5, 2.6%). Claude referenced significantly more authentic peer-reviewed scientific papers (Claude: n=111, 91.0%; Gemini: n=36, 50.7%; p<.001). The year of publication amongst all references ranged from 1988-2023, with significantly older references provided by Claude (Claude: 2008±6; Gemini: 2014±6; p<.001). Lastly, significantly more references provided by Claude were also referenced in the published NASS guidelines (Claude: n=27, 24.3%; Gemini: n=1, 2.8%; p=.04).
Conclusions:
Both generative AI tools provided responses that had acceptable alignment with NASS evidence-based guideline recommendations and offered references, though nearly a quarter of the references were inauthentic or nonscientific sources. This deficiency of legitimate scientific references does not meet standards for clinical implementation. Considering this limitation, caution should be exercised when applying the output of generative AI tools to clinical applications.
Related Concept Videos
Hindsight Biases
Stereotype Content Model
Non-equilibrium in the Cell
False Memories
One primary source of false memories is misattribution, where individuals incorrectly associate external information with...
Understanding Deception
Self-Serving Bias