召唤一个恶魔并捆绑它:LLM红色团队的接地理论
Nanna Inie1,2, Jonathan Stray3, Leon Derczynski2,4
1Paul G. Allen School of Computer Science & Engineering, Seattle, Washington, United States of America.
PloS one
|January 15, 2025
概括
这项研究将大型语言模型 (LLM) 红色团队定义为一种协作,好奇心驱动的活动,实践者故意挑起异常的LLM输出. 它揭示了在这个新领域使用的动机,策略和技术.
科学领域:
- 人工智能的人工智能
- 人与计算机的交互
- 网络安全 网络安全
背景情况:
- 大型语言模型 (LLM) 是越来越复杂的AI系统.
- 了解LLM的漏洞和故障模式对于安全部署至关重要.
- 新的人类活动正在围绕着LLM强度的故意测试而出现.
研究的目的:
- 正式定义大语言模型 (LLM) 红色团队的实践.
- 揭示攻击法学士的执业人员的动机和目标.
- 描述在LLM红色团队中使用的策略和技术.
主要方法:
- 正式的定性方法. 正式的定性方法.
- 采访数十名来自不同背景的LLM红色团队实践者.
- 基于理论的数据分析方法.
主要成果:
- 法学士红色团队被定义为寻求极限的,非恶意的,手动的,以团队为基础的活动,需要"炼金术士的心态".
- 主要动机包括内在因素,如好奇心和乐趣,以及对潜在的LLM危害的担忧.
- 确定了12种不同的攻击策略和35种特定技术的分类法.
结论:
- 这项研究提供了LLM红色团队的全面基础理论.
- 这些发现提供了对人工智能安全性和稳定性测试的人类因素的见解.
- 了解这些做法对于推动负责任的LLMs开发和部署至关重要.
相关概念视频
Language and Cognition
321
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
321
The Stanford Prison Experiment
23.0K
The famous and controversial Stanford Prison Experiment, conducted by social psychologist Philip Zimbardo and his colleagues at Stanford University, demonstrated the power of social roles, social norms, and scripts.
23.0K
Robbers Cave
14.2K
During the 1950s, the landmark Robbers Cave experiment demonstrated that when groups must compete with one another, intergroup conflict, hostility, and even violence may result. At the Oklahoman summer camp, two troops of boys—termed the Rattlers and the Eagles—took part in a week-long tournament. During this time, their negativity culminated in derogatory name-calling, fistfights, and even vandalism and destruction of property. However, this work also revealed that such tension...
14.2K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Modeling in Therapy
44
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
44


