Large reasoning models are autonomous jailbreak agents

Thilo Hagendorff1, Erik Derner2, Nuria Oliver2

  • 1University of Stuttgart, Stuttgart, Germany. thilo.hagendorff@iris.uni-stuttgart.de.

Nature Communications
|February 5, 2026
PubMed
Summary

Large reasoning models (LRMs) can now easily jailbreak AI safety features, making it simple for anyone to bypass AI security. This research highlights a critical need for improved AI alignment to prevent misuse.

Related Concept Videos

Reason and Intuition01:37

Reason and Intuition

The human brain processes information for decision-making using one of two routes: an intuitive system and a rational system (Epstein, 1994; popularized by Kahneman, 2011 as System 1 and System 2, respectively). The intuitive system is quick, impulsive, and operates with minimal effort, relying on emotions or habits to provide cues for what to do next, while the rational system is logical, analytical, deliberate, and methodical. Research in neuropsychology suggests that the...
7.5K
Reasoning01:30

Reasoning

Reasoning is the action of thinking about something in a logical, sensible way. It is integral to problem-solving, decision-making, and critical thinking. Reasoning can be inductive or deductive. Reasoning involves transforming information into conclusions, which is essential for problem-solving, decision-making, and critical thinking.
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
439
Deductive Reasoning01:16

Deductive Reasoning

Deductive reasoning, or deduction, is the type of logic used in hypothesis-based science. In deductive reasoning, the pattern of thinking moves in the opposite direction as compared to inductive reasoning, which means that it uses a general principle or law to predict specific results. From those general principles, a scientist can deduce and predict the specific results that would be valid as long as the general principles are valid.
For example, a researcher can deduce specific predictions...
69.0K
Inductive Reasoning00:59

Inductive Reasoning

Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
67.8K
Autonomic Nervous System01:22

Autonomic Nervous System

The autonomic nervous system (ANS) is a critical component of the peripheral nervous system, primarily responsible for regulating involuntary bodily functions and maintaining homeostasis. It functions in tandem with the central nervous system (CNS) to seamlessly coordinate various physiological processes without the need for conscious control.
The ANS comprises two main divisions: the sympathetic and parasympathetic divisions. These divisions function antagonistically to maintain a dynamic...
12.9K
Autonomic Nervous System: Overview01:26

Autonomic Nervous System: Overview

The human nervous system is divided into two main parts: the central nervous system (CNS) and the peripheral nervous system (PNS). The CNS is composed of the brain and spinal cord, while the PNS contains nerve cells, clusters of nerve cells, and the sensory receptors that are outside the CNS. The PNS has two types of nerve cells: sensory (afferent) and motor (efferent). Sensory cells send signals to the CNS from receptors, and motor cells carry signals from the CNS to organs, muscles, and...
7.5K