Related Experiment Video
Updated: Jan 20, 2026

Collection and Extraction of Occupational Air Samples for Analysis of Fungal DNA
Published on: May 2, 2018
CoT defender: Preemptive chain-of-thought occupation for jailbreak attack mitigation
Xiaokang Li1, Jin Liu2, Yongqiang Tang3
1Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University, Wuhan, Hubei, China; Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, Shandong, China.
CoT Defender enhances large language model (LLM) security against jailbreak attacks by using chain-of-thought analysis to block harmful content. This method improves safety without significantly impacting usability for legitimate requests.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Cybersecurity
Background:
- Large language models (LLMs) are susceptible to jailbreak attacks, compromising their safety and leading to potential abuse.
- Current defense mechanisms often fail to balance robust security with maintaining the usability of LLMs.
Purpose of the Study:
- To introduce CoT Defender, a novel defense strategy to mitigate jailbreak attacks on LLMs.
- To enhance LLM security while preserving the models' usability for benign tasks.
Main Methods:
- CoT Defender employs a chain-of-thought (CoT) analysis in the initial generated tokens to disrupt adversarial inputs.
- A two-stage training framework: Stage 1 fine-tunes for structured CoT, Stage 2 uses reinforcement learning for refinement.
- An auxiliary attacker model generates prompts, and Probabilistic Structured Output Evaluation (PSOE) provides rewards based on intent and format fidelity.
Main Results:
- Reduced average jailbreak attack success rate to below 8.0% across four tested LLMs.
- Maintained high usability, with less than a 7.0% impact on response rates for benign requests.
- Demonstrated effectiveness against six different attack methods.
Conclusions:
- CoT Defender offers an effective solution for enhancing LLM security against jailbreak attacks.
- The proposed method successfully balances robust protection with minimal impact on model usability.
- This research contributes a practical defense mechanism for safer LLM deployment.
Related Concept Videos
12:02Collection and Extraction of Occupational Air Samples for Analysis of Fungal DNA
Acid Attack on Concrete
The rate at which hydrogen...
Sulfate Attack on Concrete
Sulfates from sources like soil, groundwater, or industrial effluents...
05:50Measuring Light-Switching Behavior Using an Occupancy and Light Data Logger
Electron Transport Chains
The ETC is comprised of...
Measuring Cerebral Blood Volume in the Auditory Cortex Using Vascular Space Occupancy
