Intelligent Decision Algorithm of Target Compound Interception Based on A2C-PPO
Wei Wei Bian, Jing Wei, Kui Hua Huang, Jia Xuan Wang, Xin Lv, Wen Nan Yuan · 2021
In order to improve the timeliness of composite interception decision-making, and provide a variety of interception equipment task allocation schemes within a limited time, a proximal policy optimization (PPO) target prevention and control decision algorithm based on advantage actor critical (A2C) is proposed. Based on the in-depth analysis of A2C, temporal difference algorithm, and PPO algorithm, the architecture of A2C-PPO neural network is designed by using bidirectional long short-term memory (LSTM) network, and the training models of actor and critical network are established respectively. The situation information and entity information generated by target and intercepting equipment are taken as the input of neural network to realize the cooperative operation of intercepting equipment. Simulation results show the effectiveness of the proposed method.