Rethinking Data Integrity in Federated Learning: Are we ready?
S. R. Dixit, Parikshit Narendra Mahalle, Gitanjali R. Shinde · 2022
Federated learning is a machine learning technique that allows for the training of high-level models using data from users without the data leaving the edge device. This provides more personalized results for the user’s community. However, as federated learning is relatively new, it is vulnerable to data leakage or manipulation by attackers. The architecture of federated learning is an iterative process, so even a small amount of data manipulation can lead to significant errors. The goal of federated learning is to keep user data on the edge device. This is accomplished by sending a base machine learning model to the edge device and training it on the user’s data. The weights and biases of the trained model are then sent to the "access layer" at the server end, which collects these parameters from all the devices on the federated learning network. This means that the data is not actually sent to the server, but the model learns the weights and biases. Given the high risk of data manipulation through cyber-attacks, this paper aims to explore data integrity in transit in federated learning, including various types of cyber-attacks and activity attack modeling on the federated learning architecture. The paper concludes with a summary of the different characteristics and behavioral properties of various cyber-attacks.