Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparing the Neuropsychology Knowledge Base of Publicly Available Large Language Models
Oscar R Kronenberger1, Matthew Hutnyan1, Alyssa N Kaser1
1Department of Psychiatry, University of Texas Southwestern Medical Center, 5323 Harry Hines Blvd, Dallas, TX 75390-9044, USA.
Objective:
Assess the neuropsychology knowledge base of open-access large language models (LLMs) and inform potential applications in the field.
Method:
We obtained 600 multiple-choice practice questions from the "Be Ready for ABPP in Neuropsychology" website of the American Academy of Clinical Neuropsychology. We tested OpenAI (GPT-3.5, GPT-4, o3-mini-high) and Google (Gemini 1.0, 2.0 Flash Thinking Experimental [FTE]) models in two trials (T1: used 20-question blocks, T2: single-item re-administration of incorrect questions during T1). We compared AI-estimated to actual accuracy using a paired-samples t-test, while binomial logit generalized linear mixed-effects models (GLMMs) with a random intercept for item compared LLMs and item-level predictors (domain of practice, word count, position-in-block, and higher/lower order question type). Finally, we thematically analyzed the questions missed by the top models.
Results:
OpenAI o3-mini-high demonstrated the highest accuracy (T1:87.0%, T2:90.3%), followed by Gemini 2.0 FTE (T1:81.7%, T2:88.7%), GPT-4 (T1:74.0%, T2:85.5%), GPT-3.5 (T1:62.5%), and Gemini 1.0 (T1:52.3%). On average, LLMs overestimated their accuracy by 15.8% (87.3% vs. 71.5%, p < .001). In the GLMMs, specific LLM (p < .001) and practice domain (p = .045) were the only significant predictors of accuracy. Chain-of-thought "reasoning" models outperformed older models (p < .001) but displayed inaccuracies pertaining to aspects of neuropsychological testing and interpretation, diagnostic reasoning, clinical decision making, and neuroanatomy/neuroimaging.
Conclusions:
Chain-of-thought "reasoning" models displayed the highest accuracy, suggesting they may have utility in neuropsychology education, research, and clinical practice. However, the LLMs displayed persistent neuropsychology content area weaknesses and tended to present inaccurate information with confidence, highlighting the need for caution when interpreting LLM output.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Language and Cognition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Cognitivism
Previously dominated by behaviorism, which prioritized observable behaviors and largely ignored mental processes, psychology transformed in the 1950s. Cognitive psychologists argue that understanding how we think and process...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Concepts and Prototypes
The brain organizes this information using concepts, which are mental categories grouping linguistic data,...