人工智能和代理AI模型中的威胁和漏洞
Petar Radanliev1,2, Omar Santos3, Carsten Maple2,4
1Department of Computer Sciences, University of Oxford, Oxford, United Kingdom.
Frontiers in artificial intelligence
|March 2, 2026
概括
本研究重新定义了对代理系统的人工智能 (AI) 的对抗性脆弱性,将其从静态模型转向动态决策. 调查结果显示,漏洞在AI层之间转移,需要系统级防御来实现强大的AI安全性.
科学领域:
- 人工智能的人工智能
- 网络安全 网络安全
- 控制理论 控制理论
背景情况:
- 人工智能中的对抗性强度传统上侧重于静态模型和输入扰动.
- 代理人工智能系统通过反和闭环决策表现出动态行为,引入新的漏洞.
研究的目的:
- 重新构思人工和代理人工智能系统的对抗性脆弱性.
- 开发一个系统级的分析框架,用于AI中的对抗风险.
- 将人工智能安全的重点从基准评估转移到行为完整性和生命周期弹性.
主要方法:
- 一个符合PRISMA的系统文献审查和文献计量绘图.
- 在感知,认知和执行层面开发系统级的对抗风险分析框架.
- 来自视觉基准和大型语言模型红色团队研究的对抗性结果的综合,用于上下文化.
主要成果:
- 没有单一的防御机制可以确保代理人工智能系统的所有层的稳定性.
- 敌对的脆弱性从感知传播到政策和行动.
- 架构相似性,领域转移和反动态极大地影响了漏洞的可转移性和故障模式.
结论:
- 代理人工智能中的对抗性威胁是系统层面的风险,需要从以基准为中心的评估转变.
- 拟议的框架整合了控制理论推理和治理意识的防御设计,用于代理人工智能安全.
- 这些发现对安全关键的应用有直接影响,如自主移动,医疗成像和生物识别安全等.
相关概念视频
Non-equilibrium in the Cell
5.5K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
5.5K
Issues And Trends In Healthcare Delivery System
6.3K
The issues and trends in healthcare delivery are constantly changing. The COVID-19 pandemic is one recent issue that wreaked havoc on healthcare systems, causing a shortage of healthcare workers, high demand for medicines and supplies, and increased medical expenditure due to a lack of insurance. Other issues include rising healthcare costs and care fragmentation.
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
6.3K
Stereotype Content Model
15.6K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.6K
