Learning 2-opt Local Search from Heuristics as Expert Demonstrations

Paulo da Costa, Yingqian Zhang, Alp Akçay, Uzay Kaymak · 2021

Deep Reinforcement Learning (RL) has achieved high success in solving routing problems. However, state-of-the-art deep RL approaches require a considerable amount of data before they reach reasonable performance. This may be acceptable for small problems, but as instances grow bigger, this fact severely limits the applicability of these methods to many real-world instances. In this work, we study a setting where the agent can access data from previously handcrafted heuristics for the Traveling Salesman Problem. In our setting, the agent has access to demonstrations from 2-opt improvement policies. Our goal is to learn policies that can surpass the quality of the demonstrations while requiring fewer samples than pure RL. In this study, we propose to first learn policies with Imitation Learning (IL), leveraging a small set of demonstration data to accelerate policy learning. Afterward, we combine on policy and value approximation updates to improve performance over the expert's performance. We show that our method learns good policies in a shorter time and using less data than classical policy gradient, which does not incorporate demonstration data into RL. Moreover, in terms of solution quality, it performs similarly to other state-of-the-art deep RL approaches.

Read the paper · More papers on PaperTik