如何无辜的预训练语言模型变成隐私陷
Ruixuan Liu1, Tianhao Wang2, Yang Cao3
1Emory University, Atlanta, USA.
概括
PreCurious框架揭示了恶意预训练模型如何损害微调数据隐私. 它强调了会员推断和数据提取的风险,即使有隐私保护措施.
科学领域:
- 人工智能的人工智能
- 机器学习安全 机器学习安全
- 数据 隐私 数据 隐私 数据
背景情况:
- 标准的预训练和微调范式被广泛用于语言模型.
- 社区平台可以轻松访问预先训练的模型,但缺乏严格的验证.
- 预先训练的模型可能会对微调数据集构成隐私风险.
研究的目的:
- 引入PreCurious框架,暴露一个新的攻击面.
- 展示攻击者如何利用预先训练的模型进行隐私攻击.
- 升级隐私风险,包括会员推断和数据提取.
主要方法:
- 提出PreCurious框架来攻击微调模型.
- 操纵训练前的记忆阶段.
- 使用欺骗性的配置引导微调.
主要成果:
- PreCurious绕过了参数效率和差异性私密微调的防御.
- 该框架使隐形隐私攻击成为可能.
- 即使在严格的差异性隐私预算下,数据提取也是可能的.
结论:
- 用户必须谨慎地从不值得信赖的来源下载预训练模型.
- 常识防御和教程可能不够.
- 即使是干净的数据集也可能容易受到提取攻击.
更多相关视频
相关概念视频
Language Development
293
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
293
Social Traps
22.2K
Social traps are negative situations where people get caught in a direction or relationship that later proves to be unpleasant, with no easy way to back out of or avoid. The concept was orignally introduced by John Platt who applied psychology to Garrett Hardin's "Tragedy of the Commons", where in New England herd owners could let their cattle graze in the common ground. This situation seems like a good idea, but an individual could have an advantage. If they owned...
22.2K
Language and Cognition
303
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
303
Stereotype Threat and Self-fulfilling Prophecies
37.3K
When we hold a stereotype about a person, we have expectations that he or she will fulfill that stereotype. A self-fulfilling prophecy is an expectation held by a person that alters his or her behavior in a way that tends to make it true. When we hold stereotypes about a person, we tend to treat the person according to our expectations. This treatment can influence the person to act according to our stereotypic expectations, thus confirming our stereotypic beliefs. Research by Rosenthal and...
37.3K
Modeling in Therapy
39
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
39
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K


