Related Experiment Video
Updated: Feb 4, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Jailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations.
Large Language Models (LLMs) can be manipulated using In-Context Learning (ICL). New methods, In-Context Attack (ICA) and In-Context Defense (ICD), demonstrate how to exploit or enhance LLM safety alignment through demonstrations.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Machine Learning Security
Background:
- Large Language Models (LLMs) show great potential but are vulnerable to malicious exploitation, particularly jailbreaking attacks.
- Adversarial inputs can cause LLMs to generate harmful or unethical content, posing a significant challenge for safe deployment.
- In-Context Learning (ICL) is an effective and scalable technique for LLMs, offering a novel approach to address safety concerns.
Purpose of the Study:
- To investigate the potential of In-Context Learning (ICL) to modulate the safety alignment of Large Language Models (LLMs).
- To introduce and evaluate methods for both subverting and enhancing LLM safety using adversarial demonstrations within the ICL framework.
- To explore the theoretical and empirical implications of ICL on LLM safety and alignment.
Main Methods:
- Proposed In-Context Attack (ICA): Utilizes harmful demonstrations to compromise LLM safety alignment.
- Proposed In-Context Defense (ICD): Employs examples of refusal to harmful prompts to strengthen LLM resilience.
- Theoretical analysis and empirical validation across multiple LLM architectures, datasets, and attack baselines.
Main Results:
- Demonstrated that minimal in-context demonstrations can efficiently alter LLM safety alignment.
- ICA effectively subverts LLM safety, while ICD significantly bolsters resilience against harmful outputs.
- Both ICA and ICD showed efficacy and scalability in red-teaming evaluations and for developing safeguards.
Conclusions:
- In-Context Learning (ICL) plays a pivotal role in LLM safety and alignment, a factor previously understudied.
- The proposed ICA and ICD methods offer new tools for understanding and manipulating LLM safety.
- This research opens new avenues for improving LLM security and developing robust defenses against adversarial attacks.
More Related Videos
06:16Involving Individuals with Developmental Language Disorder and Their Parents/Carers in Research Priority Setting
Published on: June 6, 2020
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Self Within Cultural Contexts
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Impact of Social Context on Individuals