Federated Learning Empowered Violence Recognition in CCTV Footage: A YOLO and ResNet-50 Fusion Approach
Vinay Gautam, Himani Maheshwari, Raj Gaurang Tiwari, Ambuj Kumar Agarwal, Naresh Kumar Trivedi · 2024
Violence against humans in society is one major issue and cases are increasing rapidly. This creates an imbalance in society. Violence takes place suddenly in lonely places and it is difficult to handle it as there is a lack of information exchange. Although, surveillance cameras are installed at various places and the videos from these cameras can be utilized to detect violence. Several centralized deep learning(DL) and machine learning(ML) algorithms are used to address the problem, however, these methods are not suitable for protecting sensitive information. This study presented a deep learning method that uses federated learning to identify violent behaviors in CCTV video while protecting the privacy of individual users. This research uses the power of two deep learning models YOLO and RestNet-50. RestNe is trained with the data fetched with YOLO from CCTV footage available on the client side. Afterward, the trained models from clients are transferred to the server to integrate as a global model. After integration, the global model will be shared with the client for performance evaluation. It has been found that the suggested model beats another state-of-the-art deep learning model in the same situation, with an accuracy rate of 98.73%, when the data is distributed uniformly to all clients in the scenario. In research, the federated environment is set up with various hyper-performance parameters such as varying clients and communication rounds.