合成患者和面试成绩单创建器:精神健康LLM的一个重要工具
Aleyna Warner1,2, Jeffrey LeDue1,2, Yutong Cao1,2
1Department of Psychiatry, Faculty of Medicine, University of British Columbia, Vancouver, BC, Canada.
Frontiers in digital health
|September 29, 2025
概括
本研究引入了使用Llama 3.3:70B模型的合成患者和面试生成框架,以创建现实的心理健康大型语言模型 (LLM) 培训数据. 该系统确保数据隐私,同时保持人口统计准确性和多元化的语言使用.
科学领域:
- 人工智能的人工智能
- 计算语言学 计算语言学
- 心理健康技术 心理健康技术
背景情况:
- 高质量的培训数据对于开发专门的大型语言模型 (LLM) 至关重要,特别是在心理健康等敏感领域.
- 真实患者数据面临隐私和法律限制,阻碍了LLM的发展.
- 合成数据生成为克服这些数据访问挑战提供了一个潜在的解决方案.
研究的目的:
- 设计和实施一个合成的患者和面试生成框架,用于创建心理健康LLMs的培训数据.
- 确保合成数据具有丰富的上下文,人口统计准确和词汇多样性.
- 解决与使用真实患者数据相关的隐私和法律约束.
主要方法:
- 利用Llama 3.3:70B模型的两个本地运行实例:一个作为采访者,一个作为患者.
- 开发了一个可定制的问题库来构建面试成绩单.
- 采用混合方法来生成患者个人资料,将预定义的变量与LLM生成的内容相结合.
主要成果:
- 生成的采访成绩单显示了与人类对话相比的词汇多样性 (患者的Distinct-1分数中位数为0.44,采访者为0.33).
- 合成的患者个人资料显示,人口分布与现实数据没有显著差异,单词使用多样性很高 (平均Distinct-1分数为0.8).
- 该框架成功地产生了丰富的背景和现实的合成患者互动.
结论:
- 开发的框架有效地为心理健康LLMs产生高质量的合成培训数据.
- 这种方法克服了隐私和法律障碍,促进了在心理健康领域采用LLM.
- 该系统能够量身定制人口统计数据并确保数据多样性,这有助于为心理健康保健开发强大且道德的AI工具.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.2K
相关概念视频
Modeling in Therapy
381
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
381
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
