Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The Quality of Suicide-Related Stories Generated by Large Language Models
Mark Sinyor1,2, Prudence Chan1,3, Vera Yu Men4
1Department of Psychiatry, Sunnybrook Health Sciences Centre, Toronto, ON, Canada.
Abstract:
Background: Suicide-related media is known to influence suicide rates. Large language models (LLMs), a form of generative artificial intelligence (AI), are increasingly being used as a writing tool. However, the quality of LLM-generated suicide-related content has yet to be assessed. Aims: We aimed to examine suicide-related outputs from three LLMs (GPT-4, Grok, ERNIE) to characterize output quality. Methods: We provided the LLMs with 11 prompts to write different types of suicide-related content common in public media, each in five writing/content styles (broadsheet news report, tabloid news report, adult fiction, teen fiction, social media influencer). We used 3 × 2 chi square tests to compare characteristics of the outputs particularly related to adherence to responsible media guidelines for suicide reporting and overarching narratives. Results: AI outputs were generated from March 12 to July 11, 2024. A total of 147, 263, and 143 responses from GPT-4, Grok, and ERNIE were analyzed, respectively. Willingness to respond to suicide-related prompts and narrative content varied substantially across the three LLMs, with Grok producing more responses than the other two (GPT-4 = 53%, Grok = 96%, ERNIE = 52%). GPT-4 was more likely to generate emotionally supportive and antistigma messaging. Grok more frequently included harmful details such as suicide methods and romanticized portrayals. ERNIE emphasized male suicide, social support, and societal/community solutions. The proportion of outputs communicating stories of hope and recovery from a suicide crisis was both relatively low and consistent across the LLMs (16-22%). Limitations: This study examined specific prompts posed to three specific LLMs during a single epoch of time. Conclusions: LLMs produce a variety of suicide-related story content when prompted. Some stories, particularly those generated by Grok, were inconsistent with responsible media guidelines and only a minority of stories focused on hope and recovery. Engagement with AI companies to promote safer and more accurate suicide-related content is warranted as the use of LLMs continues to expand.
Related Concept Videos
Improving Translational Accuracy
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Long-term Depression
Calcium Ion Concentration Mechanism
If over time, all...
Long-term Depression
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
