Interpretable Multi-Agent Reinforcement Learning with Decision-Tree Policies *
Stephanie Milani, Zhicheng Zhang, Nicholay Topin, Zheyuan Ryan Shi, Charles Kamhoua, Evangelos E. Papalexakis, Fei Fang · 2023
Multi-agent reinforcement learning is a promising technique for solving challenging real-world problems involving potentially many interacting entities, such as air traffic control, cyber defense, and autonomous driving. However, policies trained using deep multi-agent reinforcement learning algorithms have thousands to millions of parameters and are challenging for a person to interpret and verify. Real-world risks necessitate learning interpretable policies that people can inspect and verify before deployment. At the same time, policies should perform well at the specified task and be robust to various adversaries, if applicable. This chapter introduces two algorithms, IVIPER and MAVIPER, for learning interpretable decision-tree policies in the multi-agent reinforcement learning setting. It first discusses the critical background for understanding the two algorithms, then presents a detailed explanation of IVIPER and MAVIPER. Next, the chapter includes extensive experiments to validate that MAVIPER produces high-quality decision-tree policies that can more readily coordinate. The chapter concludes by surveying related literature and commenting on avenues for future work.