Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jul 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

PsyEval: a comprehensive large language model evaluation benchmark for mental health.

Haoan Jin1, Chen Siyuan1, Dilawaier Dilixiati2

  • 1X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China.

Npj Mental Health Research
|July 9, 2026
PubMed
Summary

Related Concept Videos

Self-Evaluation Maintenance Model01:29

Self-Evaluation Maintenance Model

The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Targeting the Effector <i>AwCES</i> to Attenuate Virulence in the Postharvest Pathogen <i>Aspergillus westerdijkiae</i>.

Foods (Basel, Switzerland)·2026
Same author

MTARC1 Inactivation Remodels Lipid Droplets to Protect Against Metabolic Fatty Liver Disease.

Liver international : official journal of the International Association for the Study of the Liver·2026
Same author

Adipose tissue-specific <i>Nrf2</i> knockdown inhibits the cGAS-STING pathway to attenuate inflammation in obese mice.

Frontiers in endocrinology·2026
Same author

Highly Luminescent Zero-Dimensional (Ph<sub>3</sub>S)<sub>2</sub>HfCl<sub>6</sub>: Sb<sup>3+</sup> Hybrid Crystal Exhibiting Dual-Mode Afterglow for Information Encryption.

Inorganic chemistry·2025
Same author

DDX55 safeguards naïve T cell homeostasis by suppressing activation-promoting transposable elements.

Science immunology·2025
Same author

Audio multi-feature fusion detection for depression based on graph convolutional networks.

Annals of the New York Academy of Sciences·2025

This study introduces PsyEval, a benchmark for evaluating large language models (LLMs) in mental health tasks. Results show current LLMs struggle with accurate reasoning and appropriate responses in this sensitive domain.

Area of Science:

  • Artificial Intelligence
  • Mental Health Technology
  • Computational Psychology

Background:

  • Evaluating large language models (LLMs) in mental health is challenging due to symptom subjectivity and context dependency.
  • Existing benchmarks may not adequately capture the nuances of mental health assessment and support.

Purpose of the Study:

  • Introduce PsyEval, a novel benchmark for assessing LLMs in mental health.
  • Evaluate the performance of eleven advanced LLMs using the PsyEval benchmark.
  • Investigate the impact of prompting strategies on LLM responses in mental health contexts.

Main Methods:

  • Developed PsyEval, a benchmark focusing on knowledge, diagnosis, and emotional support in mental health.
  • Administered PsyEval to eleven diverse LLMs.

Related Experiment Videos

Last Updated: Jul 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

  • Employed various prompting strategies to test model robustness and adaptability.
  • Main Results:

    • Significant performance gaps were identified in LLMs' ability to reason accurately within mental health scenarios.
    • LLMs demonstrated limitations in providing appropriate and sensitive responses.
    • Prompting strategies showed a variable impact on model performance across different tasks.

    Conclusions:

    • Current LLMs require substantial enhancement for reliable application in mental health.
    • PsyEval offers a structured framework for future LLM development and evaluation in this domain.
    • Further research is needed to improve LLM's contextual understanding and empathetic capabilities for mental health.