Using Multi-Armed Bandit Learning for Thwarting MAC Layer Attacks in Wireless Networks
Hrishikesh Dutta, Amit Kumar Bhuyan, Subir Kumar Biswas · IEEE Transactions on Networking · 2024
This paper proposes a learning-driven approach for medium access slot allocation in the presence of malicious nodes. Learning policies are developed with the goal of defending against several forms of quasi-random slot-scheduling attack models used by the malicious nodes. The primary learning objective for the non-malicious nodes is to minimize the degradation in network performance caused by the malicious nodes. This is accomplished while minimizing the bandwidth share of the malicious nodes. These objectives are achieved using a Multi-Armed Bandit (MAB) learning architecture that allows the nodes to learn transmission schedule on-the-fly, and without the need for any central arbitrator. Two different scheduling policies are introduced: robust and reactive policies. Following the design, a detailed characterization of these policies and their use in different application-specific scenarios are presented. An analytical model of the system is developed to find the benchmark throughput for different malicious attack models. It is demonstrated that the proposed framework allows network nodes to learn close-to -benchmark slot scheduling, while thwarting attacks from the malicious nodes. The proposed architecture is validated for various mesh networks and traffic conditions in the presence of different attack models enacted by the malicious nodes.