Autonomous Drone Swarm Navigation in Complex Environments

Suleman Qamar, Maryam Qamar, Muhammad Mateen Shahbaz, Muhammad Arif Arshad, Najmus Saher Shah, Asifullah Khan · 2022 19th International Bhurban Conference on Applied Sciences and Technology (IBCAST) · 2022

Swarm intelligence is the behavior exhibited by a group of organisms performing simple actions in unison which leads to complex global behavior. This learnt behavior promotes formation of artificial swarm systems for accomplishing complex tasks and in this paper, the behavior learnt by a swarm of drones is used to accomplish multiple tasks related to complex environment navigation, obstacle avoidance and single-target tracking. In the past, manual formulation of artificial swarm systems has been attempted. The swarms usually have to be flexible entities that need to adapt to changing environments but manually crafted swarms can’t adapt to newer operating conditions thus they are not feasible. To tackle these challenges, an autonomous drone swarm navigation (ADSN) system is presented by introducing a customized architecture based on Truly Proximal Policy Optimization (TPPO) with the addition of memory cells. Furthermore, suitable 3D environments are designed for conductive learning with the addition of stability factors for drones to mimic real-life environments. Measures like Mean Cumulative Reward (MCR), Value Loss (VL), and Entropy are used to measure the performance of the presented ADSN system and other models with and without LSTM and found that incorporating memory does enhance the performance of the models although in some cases increase might not be that significant. TPPO with memory seems to have least loss and almost always performs best. Significant improvements were achieved compared to existing methodologies in convergence speed and enhanced stability of the model by customizing the architecture and hyper- parameter optimization. It is observed that Soft Actor Critic (SAC) is highly dependent on the values of the hyper-parameters and its results vary greatly with different values of hyper-parameters. The presented model has applications in maze navigation, target tracking, hover drones and other real-time scenarios.

Read the paper · More papers on PaperTik