Teaching Coordination to Selfish Learning Agents in Resource-Constrained Partially Observable Markov Games
Γεώργιος Τσαούσογλου · IEEE Transactions on Automatic Control · 2024
Of increasing relevance to engineering systems are problems that include online resource allocation to agents that feature adaptation and learning capabilities. This article considers the case where a coordinator gets to design a resource allocation mechanism (i.e., a bidding-allocation-rewards protocol) to efficiently allocate a resource to selfish agents that try to gain access by learning to communicate strategically. Toward aligning the agents' incentives with the social objective, a critical-value-based mechanism is proposed. Analytic results are presented for a simple, stylized setting, whereas simulation results for a use case with reinforcement learning agents controlling flexible loads in the smart grid demonstrate the mechanism's ability to teach coordinated behavior to the distributed learners.