Related Experiment Video
Updated: Jan 7, 2026

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
Information Bottleneck-Enhanced Reinforcement Learning for Solving Operation Research Problems
Ruozhang Xi1, Yao Ni2, Wangyu Wu3
1Krieger School of Arts and Sciences, Johns Hopkins University, Washington, DC 20001, USA.
Abstract:
Reinforcement learning (RL) has achieved remarkable success in complex decision-making tasks; however, its application to structured combinatorial optimization problems in operations research (OR) and smart manufacturing remains challenging due to high-dimensional state spaces, inefficient exploration, and unstable training dynamics. In this work, we propose Information Bottleneck-Enhanced Reinforcement Learning (IBE), a novel framework that integrates information-theoretic regularization into attention-based RL architectures to enhance both representation learning and exploration efficiency. IBE introduces two complementary objectives: (1) a state representation bottleneck, which drives the encoder to extract compact and task-relevant representations from high-dimensional sensory or operational data by minimizing redundant information; (2) a policy bottleneck, which regularizes policy optimization through an information-based exploration bonus derived from the mutual information between states and actions. Together, these mechanisms promote more robust representations, smoother policy updates, and more effective exploration in large, structured decision spaces. We evaluate IBE on representative routing and scheduling problems that commonly arise in logistics and sensor-driven manufacturing systems. Experimental results show that IBE consistently outperforms strong RL baselines, including PPO, REINFORCE, AM, and NeuOpt in both performance and stability. Comprehensive ablation studies further confirm the complementary effects of the two bottleneck components. Overall, IBE provides a principled and generalizable framework for improving RL performance in combinatorial optimization and real-world industrial decision-making under Industry 4.0 environments.
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Mathematical Modeling: Problem Solving
Gaussian Elimination: Problem Solving
Machines: Problem Solving II
Machines: Problem Solving I
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
