METREE: Max-Entropy Exploration with Random Encoding for Efficient RL with Human Preferences
Isabel Y.N Guan, Xin Liu, Gary Zhang, Estella Zhao, Zhenzhong Jia · 2023
In recent years, reinforcement learning has achieved significant advances in practical domains such as robotics. However, conveying intricate objectives to agents in reinforcement learning (RL) remains challenging, often necessitating detailed reward function design. In this study, we introduce an innovative approach, MEETRE, which integrates max-entropy exploration strategies with random encoders. This offers a streamlined and efficient solution for human-involved preference-based RL without the need for meticulously designed reward functions. Furthermore, MEETRE sidesteps the need for additional models or representation learning, leveraging the power of randomly initialized encoders for effective exploration.