Related Experiment Video
Updated: Jun 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Mitigating adversarial manipulation in LLMs: a prompt-based approach to counter Jailbreak attacks (Prompt-G)
Bhagyajit Pingua1, Deepak Murmu1, Meenakshi Kandpal1
1School of Computer Sciences, Odisha University of Technology and Research, Bhubaneswar, Odisha, India.
Prompt Guarding (Prompt-G) enhances large language model (LLM) security by detecting and filtering malicious content. This method significantly reduces jailbreak attack success rates, ensuring safer AI outputs.
Area of Science:
- Artificial Intelligence
- Cybersecurity
- Natural Language Processing
Background:
- Large language models (LLMs) offer advanced capabilities but present security risks like jailbreak attacks.
- Malicious use of LLMs can lead to misinformation and ethical concerns.
- Existing LLM vulnerabilities require robust defense mechanisms.
Purpose of the Study:
- To introduce Prompt Guarding (Prompt-G) as a novel defense against LLM jailbreak attacks.
- To evaluate the effectiveness of Prompt-G in detecting and mitigating malicious content.
- To ensure the generation of safe and accurate responses from LLMs.
Main Methods:
- Utilized vector databases and embedding techniques for text credibility assessment.
- Collected and analyzed a dataset of Self Reminder attacks.
- Integrated Prompt-G with the Llama 2 13B chat model.
Main Results:
- Prompt-G demonstrated significant reduction in jailbreak success rates.
- Effectively identified prompts causing confusion or distraction in LLMs.
- Achieved an attack success rate (ASR) of 2.08% when integrated with Llama 2 13B chat.
Conclusions:
- Prompt Guarding is an effective solution for mitigating jailbreak attacks in LLMs.
- The developed method enhances LLM safety and response accuracy.
- Real-time detection and filtering capabilities are crucial for secure AI deployment.
More Related Videos
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Stereotype Threat and Self-fulfilling Prophecies
Language and Cognition
Behavior Modification
A real-world application of operant conditioning principles is applied...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Self-Presentation: Self-Monitoring and Self-Handicapping

