Related Experiment Video
Updated: Sep 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploring AI Hallucinations of ChatGPT: Reference Accuracy and Citation Relevance of ChatGPT Models and Training
Adam Cheng1, Vikhashni Nagesh, Susan Eller
1From the KidSIM Simulation Program (A.C.), Alberta Children's Hospital, Departments of Pediatrics and Emergency Medicine, Cumming School of Medicine, University of Calgary, Calgary, Alberta, Canada; Department of Pediatrics (V.N.), Cumming School of Medicine, University of Calgary, Calgary, Alberta, Canada; Center for Immersive and Simulation-Based Learning (S.E.), Stanford School of Medicine, Stanford, CA; Departments of Pediatrics and Emergency Medicine (V.G.), Cumming School of Medicine, University of Calgary, Calgary, Alberta, Canada; and KidSIM Simulation Program (Y.L.), Alberta Children's Hospital, Calgary, Alberta, Canada.
Introduction:
Large language model-based generative AI tools, such as the Chat Generative Pre-trained Transformer (ChatGPT) platform, have been used to assist with writing academic manuscripts. Little is known about ChatGPT's ability to accurately cite relevant references in health care simulation-related scholarly manuscripts. In this study, we sought to: (1) determine the reference accuracy and citation relevance among health care simulation debriefing articles generated by 2 different models of ChatGPT and (2) determine if ChatGPT models can be trained with specific prompts to improve reference accuracy and citation relevance.
Methods:
The ChatGPT-4 and ChatGPT o1 models were asked to generate scholarly articles with appropriate references based upon three different article titles about health care simulation debriefing. Five articles with references were generated for each article title-3 ChatGPT-4 training conditions and 2 ChatGPT o1 training conditions. Each article was assessed independently by 2 blinded reviewers for reference accuracy and citation relevance.
Results:
Fifteen articles were generated in total: 9 articles by ChatGPT-4 and 6 articles by ChatGPT o1. A total of 60.4% of the 303 references generated across 5 training conditions were classified as accurate, with no significant difference in reference accuracy between the 5 conditions. A total of 22.2% of the 451 citations were classified as highly relevant, with no significant difference in citation relevance across the 5 conditions.
Conclusions:
Among debriefing articles generated by ChatGPT-4 and ChatGPT o1, both ChatGPT models are unreliable with respect to reference accuracy and citation relevance. Reference accuracy and citation relevance for debriefing articles do not improve even with some degree of training built into ChatGPT prompts.
Related Concept Videos
Improving Translational Accuracy
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
The Representativeness Heuristic
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Hindsight Biases
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...

