Apprentissage par renforcement profond à travers l'apprentissage par imitation et l'apprentissage par curriculum : application à la planification des pompes dans les réseaux de distribution d'eau
Henrique Donãncio Nunes Rodrigues · HAL (Le Centre pour la Communication Scientifique Directe) · 2023
Water distribution systems must be carefully monitored to ensure water supply, prevent losses, and protect system assets. In addition, optimizing pump scheduling in such systems can lead to electricity savings without incurring additional costs. However, these systems are usually controlled by Supervisory Control and Data Acquisition (SCADA) systems, which can be expensive and complex. The Internet of Things (IoT) opens a new path in interoperability for data-driven approaches to controlling these systems, where sensors could collect near-real-time data to improve decision-making. Furthermore, approaches such as Deep Q-Networks provide a scalable data-driven decision-making architecture to handle real-world control problems.Despite the scalability of approaches based on Deep Q-Networks, collecting experiences through real-world interactions can be costly or even infeasible. This difficulty stems from the fact that real-world applications can be risk-sensitive, and building simulators of these systems can be a complex task. Even if there is no such limitation, exploring an environment with high dimensional state space can be costly. To overcome exploration, we present an Imitation Learning approach where the reward function is augmented to encourage policy convergence to known states, given a demonstration distribution from an expert policy. Later, we present a Curriculum Learning approach combined with Transfer Learning/Policy Distillation. We decompose a target task into more straightforward ones along a curriculum, transferring the knowledge in earlier steps to later ones to leverage the learning process.