PsychiatryBench: a multi-task benchmark for LLMs in psychiatry.

Aya E Fouda1, Abdelrahman A Hassan1, Radwa J Hanafy1,2

  • 1Compumacy for Artificial Intelligence Solutions, Cairo, Egypt.

NPJ Digital Medicine
|April 14, 2026
PubMed
Summary

PsychiatryBench, a new benchmark for large language models (LLMs), uses expert-validated psychiatric texts for evaluation. It reveals LLMs struggle with clinical consistency and safety in complex mental health tasks.

Related Concept Videos