A Fault-tolerant and Cost-efficient Workflow Scheduling Approach Based on Deep Reinforcement Learning for IT Operation and Maintenance
Yunsong Xiang, Xuemei Yang, Yan Sun, Hong Luo · 2023
With the promotion of cloud computing, a large number of hardware and software systems in the cloud bring massive and complex operation and maintenance (O&M) work. To ensure the O&M efficiency of IT infrastructures, it is necessary to implement automatic and reliable scheduling for the directed acyclic graph (DAG) workflow which is composed of multiple O&M tasks. Considering the changing status of networks and machines in the cloud and the position constraints that some tasks must be executed on the specified machines in some O&M scenarios, we propose a novel workflow scheduling approach based on Deep Reinforcement Learning (DRL) to minimize the workflow execution makespan and implement the fault tolerance with the position constraints of tasks execution. In our proposal, we first design a fault-tolerant mechanism according to the reliability requirement and the probability distributions of the machine failure parameters with consideration of different failure rates in the heterogeneous environment. Then, we employ proximal policy optimization (PPO) to optimize the task scheduling strategy and ensure the strategy to satisfy the position constraints of tasks execution by action masking in proximal policy optimization. The experimental results show that our proposal can effectively reduce the makespan of the fault-tolerant workflow on the premise of 99.9% reliability.