Related Experiment Video
Updated: Jul 16, 2026

Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
Reference Hallucination, Citation Reliability, and Readability of Large Language Models in Anatomy-Related Question
Mehmet Ülkir1, Bahattin Paslı1
1Department of Anatomy, Faculty of Medicine, Hacettepe University, Ankara, Türkiye.
Abstract:
Large language models (LLMs) are increasingly used in medical education and academic writing. However, concerns remain regarding reference hallucination, citation, and the reliability of LLM-generated content. This study aimed to evaluate the performance of ChatGPT 5.2, Gemini 3 Pro, and DeepSeek V3.2 in generating anatomy-related responses by assessing bibliographic reference accuracy, citation content consistency, and the readability of LLM-generated content. A total of 120 open-ended anatomy questions covering six anatomical categories (neuroanatomy, musculoskeletal, respiratory and circulatory, gastrointestinal, urogenital and endocrine, head and neck) were submitted to each model. Individual citation components, including author names, article titles, journal names, publication details, and PMIDs, were verified against indexed sources. Citation content consistency was evaluated using a three-point Likert scale. Readability was assessed using the Flesch Reading Ease score, Flesch-Kincaid Grade Level, Coleman-Liau, and Simple Measure of Gobbledygook indices. A total of 1800 references were analyzed. ChatGPT 5.2 demonstrated the lowest hallucination rate (23.2%), whereas Gemini 3 Pro and DeepSeek V3.2 exhibited substantially higher hallucination rates (45.8% and 47.5%, respectively). DeepSeek V3.2 achieved the highest accuracy for several individual bibliographic components, including author names, article titles, volumes, issues, pages, and journal names. PMID accuracy remained limited across all models, ranging from 25.1% to 57.6%. Citation content consistency differed significantly among the models (p < 0.001), with ChatGPT 5.2 demonstrating the highest proportion of fully supported citations (67.2%), compared with Gemini 3 Pro (42.5%) and DeepSeek V3.2 (41.0%). Citation accuracy differed significantly across most anatomical subcategories, with the greatest intermodel discrepancy observed in head and neck anatomy. Readability analyses indicated that the generated responses generally required college-level reading proficiency. Although LLMs can generate plausible anatomy-related responses, substantial limitations remain regarding reference accuracy, hallucination, and citation reliability. Human verification remains essential before incorporating LLM-generated references into academic or educational materials.
