Related Experiment Video
Updated: Feb 28, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Snake Oil or Panacea? How to Misuse AI in Scientific Inquiries of the Human Mind
René Schlegelmilch1, Lenard Dome2,3
1Department of Psychology, Faculty of Human and Health Sciences, University of Bremen, 28359 Bremen, Germany.
Abstract:
Large language models (LLMs) are increasingly used to predict human behavior from plain-text descriptions of experimental tasks that range from judging disease severity to consequential medical decisions. While these methods promise quick insights without complex psychological theories, we reveal a critical flaw: they often latch onto accidental patterns in the data that seem predictive but collapse when faced with novel experimental conditions. Testing across multiple behavioral studies, we show these models can generate wildly inaccurate predictions, sometimes even reversing true relationships, when applied beyond their training context. Standard validation techniques miss this flaw, creating false confidence in their reliability. We introduce a simple diagnostic tool to spot these failures and urge researchers to prioritize theoretical grounding over statistical convenience. Without this, LLM-driven behavioral predictions risk being scientifically meaningless, despite impressive initial results.
Related Concept Videos
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
Non-equilibrium in the Cell
False Memories
One primary source of false memories is misattribution, where individuals incorrectly associate external information...
Reason and Intuition
Cause and Effect
Introspection
