Related Experiment Video
Updated: Sep 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Systemic Failure Modes of Large Language Models in Geriatric Psychiatric Assessment: A Scoping Review
1Department of Psychiatry (Y-SL), Taipei Veterans General Hospital, Taipei, Taiwan.
Background:
Large language models (LLMs) show strong performance on medical examinations, yet their reliability in real-world psychogeriatric assessment is uncertain, where delirium, dementia, and late-life depression often coexist amid frailty, polypharmacy, and long clinical histories. We aimed to map empirically evaluated LLM failure modes most relevant to geriatric psychiatric assessment and risk management.
Methods:
We conducted a systematic scoping review of empirical evaluations published between January 1, 2023, and December 31, 2025. We searched biomedical and technical sources and included studies that tested LLMs on clinically relevant tasks, including diagnostic reasoning, longitudinal history integration, safety and medication reasoning, and bias-related outcomes. Evidence was charted and synthesized qualitatively to derive a pragmatic taxonomy of recurrent failure patterns.
Results:
Forty-seven studies met inclusion criteria. Four convergent failure modes emerged: (1) diagnostic instability in longitudinal reasoning, including degraded retrieval within long contexts and loss of mid-history cues; (2) adversarial vulnerability, with elevated hallucination under misleading inputs and flawed rationales despite correct final answers; (3) multimodal temporal sparsity in video-based models that can omit brief safety-critical events such as falls; and (4) systemic ageism and value misalignment that can distort clinical narratives, risk estimation, and care recommendations.
Conclusions:
Current LLM architectures remain insufficient for autonomous geriatric psychiatric assessment. Near-term use should be limited to low-risk support functions with human oversight, structured cognitive forcing safeguards, and evaluation protocols that stress-test longitudinal complexity, adversarial conditions, multimodal safety events, and bias.
Related Concept Videos
Drug Dosing: Geriatric Patients
Alzheimer Disease l: Introduction
Dementia l: Introduction
Psychosis: Goals of Pharmacotherapy
Pharmacodynamics in Geriatric Patients: Effects of Age