Toward Multi-Agent Coordination in IoT via Prompt Pool-based Continual Reinforcement Learning
Chenhang Xu, Jia Wang, Xiaohui Zhu, Yong Yue, Jun Qi, Jieming Ma · 2024
The Internet of Things (IoT) represents a complex, dynamic environment where edge devices continuously optimize their policies to address a continual stream of tasks. Previous studies have typically relied on a rehearsal buffer containing data from past tasks or a known task identity to mitigate catastrophic forgetting. Our research, Prompt Pool-based Continual Reinforcement Learning (PPCRL), aims to create a more efficient memory system by expanding a single prompt into a prompt pool, allowing agents to automatically select a set of relevant prompts without needing task identity knowledge. Similar to prompt-based learning techniques, our approach utilizes a small trainable prompt pool to guide pre-trained models through sequential task learning systematically. This allows us to optimize prompts for guiding model predictions and effectively manage both shared and task-specific knowledge while maintaining model generalization. We conducted experiments on two multi-agent benchmarks where traditional methods suffer from significant performance degradation. In contrast, PPCRL demonstrates the capability to outperform baselines and exhibits high generalization ability.