FUZZ-PPO: Fuzzy-Proximal Policy Optimisation for Enhanced Robotic Control in Dynamic Environments Using Meta-Reinforcement Learning

Baljinder Sanghera, Kartikeya Walia · 2025

This paper introduces Fuzzy-Proximal Policy Optimization (Fuzz-PPO), a unique hybrid control framework that improves robotic control in dynamic and unpredictable situations by utilising Meta-Reinforcement Learning (Meta-RL). Typically, traditional robotic control systems require significant training data and struggle to achieve fast convergence and stability, especially under dynamic conditions. To overcome these problems, the suggested Fuzz-PPO model makes use of Meta-RL's generalisation capabilities, which enable robots to seamlessly adjust to shifting situations without requiring retraining. The use of a fuzzy logic controller also strengthens the original policy's resilience, enhancing its overall functionality. Utilising the CartPole-v1 environment, the hybrid Fuzz-PPO with Meta-RL was evaluated in a dynamic control simulation, where the goal was to stabilise the inverted pendulum whilst adapting to the system's dynamic changes. The findings show that the suggested method performs noticeably better than traditional Reinforcement Learning (RL) and PPO-based methods. Specifically, Fuzz-PPO achieved 30% faster convergence compared to traditional PPO, with a 20% improvement in stability during the learning process. Additionally, the model demonstrated 15% higher balancing control accuracy, reducing the occurrence of instability in dynamic environments. These findings support the claim that the hybrid Fuzz-PPO with Meta-RL is more resilient to fluctuations and better equipped to handle uncertain conditions. This paper contributes to the advancement of robotic control by providing a highly flexible, efficient, and real-time decision-making model that excels in managing complex and dynamic environments.

Read the paper · More papers on PaperTik