Related Experiment Video
Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Synthetic patient and interview transcript creator: an essential tool for LLMs in mental health
Aleyna Warner1,2, Jeffrey LeDue1,2, Yutong Cao1,2
1Department of Psychiatry, Faculty of Medicine, University of British Columbia, Vancouver, BC, Canada.
This study introduces a synthetic patient and interview generation framework using Llama 3.3:70B models to create realistic training data for mental health large language models (LLMs). The system ensures data privacy while maintaining demographic accuracy and diverse language use.
Area of Science:
- Artificial Intelligence
- Computational Linguistics
- Mental Health Technology
Background:
- High-quality training data is crucial for developing specialized large language models (LLMs), particularly in sensitive fields like mental health.
- Real patient data faces privacy and legal restrictions, hindering LLM development.
- Synthetic data generation offers a potential solution to overcome these data access challenges.
Purpose of the Study:
- To design and implement a synthetic patient and interview generation framework for creating training data for mental health LLMs.
- To ensure the synthetic data is contextually rich, demographically accurate, and lexically diverse.
- To address privacy and legal constraints associated with using real patient data.
Main Methods:
- Utilized two locally run instances of Llama 3.3:70B models: one as an interviewer and one as a patient.
- Developed a customizable question bank to structure interview transcripts.
- Employed a hybrid approach for patient profile generation, combining predefined variables with LLM-generated content.
Main Results:
- Generated interview transcripts exhibited lexical diversity comparable to human conversation (median Distinct-1 scores of 0.44 for patient, 0.33 for interviewer).
- Synthetic patient profiles showed demographic distributions not significantly different from real-world data and high word usage diversity (average Distinct-1 score of 0.8).
- The framework successfully produced contextually rich and realistic synthetic patient interactions.
Conclusions:
- The developed framework effectively generates high-quality synthetic training data for mental health LLMs.
- This approach overcomes privacy and legal barriers, facilitating the adoption of LLMs in mental health.
- The system's ability to tailor demographics and ensure data diversity supports the development of robust and ethical AI tools for mental healthcare.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Improving Translational Accuracy
Improving Translational Accuracy