AIPsychoBench:了解LLM和人类之间的心理测量差异
Wei Xie1, Zhenhua Wang1, Shuoyoucheng Ma1
1College of Computer Science and Technology, National University of Defense Technology.
Topics in cognitive science
|March 9, 2026
概括
AIPsychoBench是评估大型语言模型 (LLM) 心理属性的新基准. 它提高了响应率并减少了偏见,为LLM心理测量提供了对语言影响的见解.
科学领域:
- 人工智能的人工智能
- 计算心理学 计算心理学
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 显示出类似人类的智能,但存在不可解释性问题,引发了可靠性问题.
- 目前对LLM的心理测量评估由于LLM和人类之间的根本差异而失败,导致高排斥率.
- 目前的方法不考虑语言差异,限制了对不同语言的LLM心理属性的评估.
研究的目的:
- 介绍AIPsychoBench,这是一个专门的基准,旨在准确评估LLMs的心理特性.
- 通过解决以人为中心的心理尺度的局限性,提高法学士评估的可解释性和可靠性.
- 研究语言多样性对LLM心理测量的影响.
主要方法:
- 开发了一个轻量级的角色扮演提示以绕过LLM对齐,提高有效响应率.
- 将新提示方法与传统的越狱提示符的偏差水平进行比较.
- 在七种语言中评估了112个心理测量子类别,以测量与英语相对的得分偏差.
主要成果:
- 角色扮演提示显著改善了平均有效响应率,从70.12%提高到90.40%.
- 与越狱提示 (9.8%的积极,6.9%的负面) 相比,新方法的平均偏差要低得多 (3.3%的积极,2.1%的负面).
- 七种语言的分数偏差在43个子类别中从5%到20.2%不等,表明对LLM心理测量的语言影响很大.
结论:
- AIPsychoBench提供了一种更可靠,更有效的方法来评估LLM的心理特性.
- 该基准表明偏差减少和响应率更高,提高了LLM的解释性.
- 这项研究提供了第一个关于语言变异影响LLM心理测量评估的全面证据.
相关概念视频
Lateralization
1.2K
Brain lateralization refers to the division of mental processes and functions between the two hemispheres of the brain, a phenomenon that optimizes neural efficiency and underpins complex abilities in humans. This specialization allows each hemisphere to perform tasks where it has a comparative advantage, facilitating more refined cognitive capabilities across different domains.
1.2K
Self-Report Tests of Personality
1.1K
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
1.1K


