基于强化学习的多重欺骗资源部署策略,用于网络威胁缓解和缓解网络威胁
Changsong Li1,2,3, Ning Zhao4, Hao Wu5,6,7,8
1State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University, Beijing, 100044, China.
Scientific reports
|May 14, 2025
概括
本研究介绍了一种强化学习算法,用于对高级持久威胁 (APT) 部署欺骗资源. 该方法优化了防御效率和成本,在保护关键资产方面实现了97.09%的成功率.
科学领域:
- 网络安全 网络安全
- 人工智能的人工智能
- 网络防御 网络防御 网络防御
背景情况:
- 高级持久威胁 (APT) 对信息系统安全构成重大挑战.
- 主动防御方法,如使用 honeypots 的欺骗防御 (DD),是常见的,但面临着资源限制.
- 有效地部署欺骗资源对于减轻APT风险至关重要.
研究的目的:
- 提出一种基于强化学习的算法,用于生成多种类型的欺骗资源部署策略.
- 解决在APT网络侦察阶段在资源有限的环境中部署欺骗资源的挑战.
- 在APT缓解中平衡防御效率和防御成本.
主要方法:
- 开发了一个强化学习算法,用于生成欺骗资源部署策略.
- 分析网络资产和攻击过程,以告知策略生成.
- 作为关键优化参数,平衡国防有效性和国防成本.
主要成果:
- 实现了97.09%的防守成功概率,超过了基线方法.
- 将目标资产的攻击概率降低了至少10.34%.
- 证明了算法融合,稳定性和降低国防成本的效率.
结论:
- 拟议的算法有效地提高了对APT的防御效率.
- 它提供了一个可行的解决方案,用于优化在受约束系统中的欺骗资源部署.
- 这种方法成功地降低了国防成本,同时保持了高安全效率.
相关概念视频
Reinforcement
167
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
167
Modeling in Therapy
37
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
37
Law of Effect
1.3K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.3K
Role of Shaping in Operant Conditioning
232
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
232
Operant Conditioning Intervention
30
Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...
In operant conditioning, behaviors that are...
30
Primary and Secondary Reinforcers
144
In psychology, reinforcement is a key concept in behavior modification. B.F. Skinner demonstrated this with his experiments involving rats in what is known as a Skinner box. The rats learned to press a lever to receive food, a primary reinforcer that fulfilled their innate need for nourishment.
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
144


