Related Experiment Videos
Collaborative-adversarial jailbreaking: A propagation-aware attack framework for multi-agent code generation systems
Zhaoyang Qu1, Mingyang Geng2, Yunxin Mao1
1College of Computer Science and Technology, National University of Defense Technology, Changsha, Hunan, 410073, China.
Abstract:
The rapid adoption of multi-agent frameworks for automated code generation has significantly enhanced software development efficiency, yet simultaneously introduced critical security challenges that remain largely unexplored. While extensive research has investigated jailbreaking vulnerabilities in single-agent large language models, existing studies have overlooked the unique security risks arising from collaborative dynamics in multi-agent systems, where distributed decision-making and social interactions may amplify rather than mitigate adversarial threats. To address this gap, we propose the first comprehensive security assessment framework for multi-agent code generation, introducing Implicit Multi-agent Attack (IMA), a novel jailbreaking strategy that exploits social engineering and collaborative reinforcement within agent networks. Our evaluation encompasses four prominent frameworks (MetaGPT, CrewAI, AutoGen, and ChatDev) using the established RMCBench benchmark (Resistance to Malicious Code Benchmark) with 282 malicious code generation tasks across text-to-code, function-level, and block-level completion scenarios. Compared to traditional explicit attacks and single-agent baselines, IMA demonstrates substantially higher effectiveness, achieving an average attack success rate of 89.01% and revealing collaborative harm amplification factors up to 114.9%. The results expose fundamental vulnerabilities in current multi-agent architectures, with defense mechanisms showing alarmingly low detection rates below 30%. Crucially, by isolating the final coding agent via a direct jailbreak baseline (SADJ), we demonstrate that multi-agent collaboration itself amplifies attack success by an average of 11.4% (CADA), confirming that the vulnerability lies in the architecture rather than the underlying model alone.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
Catalysis
Catalysis
Statically Indeterminate Problem Solving
A Single-Component System
Synthetic Biology
Golden rice
Golden rice is a genetically modified...