Information Bottleneck-Enhanced Reinforcement Learning for Solving Operation Research Problems
Ruozhang Xi, Yao Ni, Wangyu Wu · Sensors · 2025
Reinforcement learning (RL) has achieved remarkable success in complex decision-making tasks; however, its application to structured combinatorial optimization problems in operations research (OR) and smart manufacturing remains challenging due to high-dimensional state spaces, inefficient exploration, and unstable training dynamics. In this work, we propose Information Bottleneck-Enhanced Reinforcement Learning (IBE), a novel framework that integrates information-theoretic regularization into attention-based RL architectures to enhance both representation learning and exploration efficiency. IBE introduces two complementary objectives: (1) a state representation bottleneck, which drives the encoder to extract compact and task-relevant representations from high-dimensional sensory or operational data by minimizing redundant information; (2) a policy bottleneck, which regularizes policy optimization through an information-based exploration bonus derived from the mutual information between states and actions. Together, these mechanisms promote more robust representations, smoother policy updates, and more effective exploration in large, structured decision spaces. We evaluate IBE on representative routing and scheduling problems that commonly arise in logistics and sensor-driven manufacturing systems. Experimental results show that IBE consistently outperforms strong RL baselines, including PPO, REINFORCE, AM, and NeuOpt in both performance and stability. Comprehensive ablation studies further confirm the complementary effects of the two bottleneck components. Overall, IBE provides a principled and generalizable framework for improving RL performance in combinatorial optimization and real-world industrial decision-making under Industry 4.0 environments.