RIB-Guard: A Risk-Aware Information Bottleneck Defense for Black-Box Large Language Models

Muen Cai1, Yuan Shen2, Xiong Luo3

  • 1School of Computer Science and Engineering, University of Electronic Science and Technology of China, No. 2006, Xiyuan Ave, Chengdu 611731, China.

Summary

This study introduces RIB-Guard, a novel defense against large language model (LLM) jailbreaks in black-box settings. RIB-Guard enhances LLM security by learning a token-level masking policy for improved prompt protection.