Efficient Retraining for Continuous Operability Through Strategic Directives

Gentoku Nakasone, Yoshinari Motokawa, Yuki Miyashita, Toshiharu Sugawara · 2024

We introduce a method to ensure the controllability of agents in pretrained networks, anticipating changes in managerial requirements and environmental conditions. Advances in multi-agent deep reinforcement learning (MADRL) have fostered sophisticated cooperative behaviors in multi-agent systems in which agents share and execute various complex tasks. However, retraining MADRL systems to adapt to new conditions is costly, because many agents are involved. Our approach introduces several types of directives, termed destination channels (DCs), which allow agents to experience diverse coordination patterns during training without detailed instructions. When changes occur, the system manager assigns appropriate DCs to each agent to facilitate adaptation and maintain a continuous operation. We conducted experiments using object-collection games to evaluate our proposed method by comparing the number of objects recovered and the learning speed of our method with those of existing methods, an implicit quantile network (IQN), and a traditional deep Q-network (DQN). The experimental results demonstrate that agents using this method adapt more swiftly to environmental changes during retraining than baseline methods, sustaining performance without significant degradation.

Read the paper · More papers on PaperTik