Related Experiment Videos
D3O-IIoT: deep reinforcement learning-driven dynamic deception orchestration for industrial IoT security
Usman Wushishi1, Altaf Hussain2, Muhammad Imran Khalid1
1School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, 400065, China.
Scientific Reports
|December 21, 2025
Summary
Industrial Internet of Things (IIoT) cybersecurity is enhanced by D3O-IIoT, a reinforcement learning model that dynamically orchestrates deception techniques. This adaptive defense system significantly improves attack mitigation while minimizing false alarms in resource-constrained industrial environments.
Area of Science:
- Cybersecurity
- Artificial Intelligence
- Industrial Internet of Things (IIoT)
Background:
- Industrial Internet of Things (IIoT) systems face increasing cyber threats exploiting resource limitations and operational vulnerabilities.
- Existing intrusion detection systems lack dynamic adaptation to evolving attack vectors, relying on static or passive defenses.
Purpose of the Study:
- To present D3O-IIoT, a novel reinforcement learning model for dynamic deception orchestration in IIoT cybersecurity.
- To dynamically coordinate deception techniques like honeypots, moving target defense, fake telemetry, and node isolation based on real-time threat monitoring.
Main Methods:
- Formulated the defense problem as a Markov Decision Process.
- Employed a Dueling Deep Q-Network agent to maximize a multi-objective reward balancing attack mitigation, deception engagement, false positive control, and resource cost.
- Validated the model on three IIoT datasets (CIC-IIoT2025, WUSTL-IIoT2021, TON-IoT).
Main Results:
- Achieved a 13.7% attack mitigation rate with a 0.3% false alarm rate, significantly outperforming baseline methods (293-767% improvement).
- Demonstrated strong generalization capabilities with high retention rates across different datasets (97.7% on TON-IoT, 77.8% on WUSTL-IIoT).
- Identified false positive control as the most critical reward component and showed tunability through risk threshold adjustments.
Conclusions:
- D3O-IIoT offers a dynamic, learning-based deception orchestration for IIoT cybersecurity, moving beyond fixed rule-based defenses.
- The model effectively balances multiple practical objectives, including attack mitigation and resource cost, under constrained conditions.
- Real-time deployment is feasible with low latency (2.07ms), favoring isolation for confirmed threats and honeypots for reconnaissance.
Related Concept Videos
Understanding Deception
142
Deception is a pervasive aspect of human communication. Empirical studies have shown that most individuals engage in some form of deceit on a daily basis, with approximately 20% of social exchanges involving deceptive elements. Lying follows a developmental trajectory, peaking during adolescence and declining with age, possibly due to the maturation of cognitive control and social accountability.Cognitive and Social Factors in Deception DetectionDespite its prevalence, accurately detecting...
142
Masking and Demasking Agents
3.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
3.4K
Observational Learning
791
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
791
Reinforcement
786
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
786
Cognitive Learning
970
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
970
Distribution Reliability and Automation
474
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
474