Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 9, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
07:14

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models

Published on: December 23, 2025

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation.

Kate H Bentley1,2, Luca Belli1,3, Adam M Chekroud1,4

  • 1Spring Health, New York, NY, United States.

JMIR AI
|June 8, 2026
PubMed
Summary

Related Concept Videos

Survey Safety01:28

Survey Safety

Surveying near highways, rough terrain, or power lines involves significant risks. Working along highways is particularly dangerous and requires the use of warning signs and flagmen. It is safest to avoid working directly on roads and use offsets whenever possible. When highway work is unavoidable, it must follow all safety guidelines. Surveyors should wear bright clothing, such as orange reflective vests, to ensure visibility to motorists, coworkers, and hunters. In construction zones, wearing...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A large-scale evaluation of provider-patient matching in an employer-sponsored mental health program.

Npj mental health research·2026
Same author

Using smartphone surveys to predict next-week suicide attempts.

Journal of psychopathology and clinical science·2026
Same author

Emotion Reactivity Moderates the Association Between Momentary Negative Affect and Suicidal Thinking.

Brain and behavior·2026
Same author

User Experience and Early Clinical Outcomes of a Mental Wellness Chatbot for Depression and Anxiety: Pilot Evaluation Mixed Methods Study.

JMIR formative research·2026
Same author

Momentary emotion-related distress, negative emotions, and suicidal ideation: Findings from an RCT among adults in inpatient psychiatric care.

Journal of affective disorders·2026
Same author

SAFEGUARD: Transforming Military Suicide Prevention Through Predictive Analytics and Targeted Interventions.

Current psychiatry reports·2026

The Validation of Ethical and Responsible AI in Mental Health (VERA-MH) benchmark shows strong agreement between AI chatbot safety evaluations and expert clinicians, particularly for suicide risk detection and response.

Area of Science:

  • Artificial Intelligence in Mental Health
  • Clinical Evaluation of AI Safety

Background:

  • Generative AI chatbots are increasingly used for psychological support, raising critical safety concerns.
  • A validated, automated benchmark is needed to assess AI chatbot safety, especially for suicide risk.
  • The Validation of Ethical and Responsible AI in Mental Health (VERA-MH) evaluation was developed to address this need.

Purpose of the Study:

  • To validate the VERA-MH safety evaluation by comparing its automated assessments with expert clinician ratings.
  • To examine the alignment between AI chatbot safety evaluations and human clinical judgment for suicide risk detection and response.

Main Methods:

  • Simulated conversations with varying suicide risk levels and disclosure styles were created.
  • Licensed mental health clinicians rated conversations using the VERA-MH rubric.
Keywords:
artificial intelligencebenchmarkchatbotmental healthsafetysuicide

Related Experiment Videos

Last Updated: Jun 9, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
07:14

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models

Published on: December 23, 2025

  • An AI-based evaluator (LLM judge) also rated conversations using the same rubric.
  • Alignment was assessed between clinicians, between clinician consensus and the LLM judge, and across different LLMs.
  • Main Results:

    • High inter-rater reliability (IRR: 0.77) was found among clinicians, establishing a reliable consensus.
    • The LLM judge demonstrated strong alignment with clinical consensus (IRR: 0.81).
    • AI safety ratings were stable across different LLMs and evaluation iterations.
    • Clinician ratings on the realism of simulated users and their disclosed risk/disclosure styles were mixed.

    Conclusions:

    • The VERA-MH evaluation is a reliable, open-source, automated tool for assessing AI chatbot safety in suicide risk.
    • Ensuring AI safety is crucial for realizing the potential mental health benefits of AI chatbots.
    • Future work should focus on external validation of evolving VERA-MH versions, generalizability, robustness, and expanding to other AI safety areas in mental health.