Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large
Jake Linardon1, Hannah K Jarman1, Zoe McClure1
1School of Psychology, Faculty of Health, Deakin University, Geelong, Victoria, Australia.
Large language models (LLMs) like GPT-4o frequently fabricate citations in mental health reviews, with nearly two-thirds of references being inaccurate. Reliability varies by topic visibility and prompt specificity, emphasizing the need for human verification.
Area of Science:
- Artificial Intelligence in Mental Health Research
- Natural Language Processing Applications
- Scholarly Communication Integrity
Background:
- Large language models (LLMs) are increasingly used in mental health research for efficiency.
- LLMs can generate plausible but fabricated content, including non-existent bibliographic citations.
- Previous research on citation fabrication lacks a systematic analysis across mental health topics with varying public visibility and scientific maturity.
Purpose of the Study:
- To assess citation fabrication and bibliographic errors in GPT-4o generated literature reviews on mental health topics.
- To investigate how public familiarity and scientific maturity of mental health disorders influence citation accuracy.
- To determine if prompt specificity (general vs. specialized) impacts fabrication and accuracy rates.
Main Methods:
- GPT-4o generated six literature reviews (~2000 words, ≥20 citations) on major depressive disorder (high visibility), binge eating disorder (moderate), and body dysmorphic disorder (low).
- Reviews were conducted at general (symptoms, impacts, treatments) and specialized (digital interventions) levels.
- A total of 176 citations were extracted and verified using multiple academic databases; fabrication and accuracy rates were compared using chi-square tests.
Main Results:
- GPT-4o generated 176 citations, with 35 (19.9%) being fabricated and 64 (45.4%) of the real citations containing errors.
- Fabrication rates significantly differed by disorder (P=.001), being higher for binge eating disorder (28%) and body dysmorphic disorder (29%) compared to major depressive disorder (6%).
- Specialized reviews showed higher fabrication rates for binge eating disorder (46% vs. 17%; P=.01), and accuracy varied by disorder and review type.
Conclusions:
- Citation fabrication and bibliographic errors are prevalent in GPT-4o outputs, affecting nearly two-thirds of references.
- The reliability of LLM-generated citations is influenced by disorder familiarity and prompt specificity, posing greater risks for less visible or specialized topics.
- Rigorous human verification, careful prompt design, and institutional safeguards are crucial for maintaining research integrity with LLM integration.
More Related Videos
Related Concept Videos
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Stereotype Content Model
Fundamental Attribution Error
Hindsight Biases
Stereotype Threat and Self-fulfilling Prophecies
The Anchoring-and-Adjustment Heuristic

