PsyEval: a comprehensive large language model evaluation benchmark for mental health.

Haoan Jin1, Chen Siyuan1, Dilawaier Dilixiati2

  • 1X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China.

Summary

This study introduces PsyEval, a benchmark for evaluating large language models (LLMs) in mental health tasks. Results show current LLMs struggle with accurate reasoning and appropriate responses in this sensitive domain.