AlertMind: A Multimodal Approach for Detecting Menace using Transformers

Rahul Mallya, Pratham Deepak Rao, Purvik S Nukal, Rahul Ranganath, V R Badri Prasad · 2024

Amidst the unprecedented surge in online video content, the need for effective automated violence detection systems has become paramount. This is particularly crucial given that exposure to violence can significantly impact the mental health of individuals watching such content. In this study, an audio-visual guided framework for violence detection has been proposed that utilizes audio and video inputs to accurately identify violence in a large variety of videos and alert the user in real-time. The aim of this study is to capture input from a wide range of sources and process the video and audio input using computer vision and signal processing techniques respectively. The features of these modalities are leveraged to classify events as Violent or Non-violent and subsequently, identify the specific type of violence using deep learning transformer models. This framework operates in real-time and can be scaled to monitor social media or surveillance systems with multiple cameras and microphones simultaneously, making it ideal for real-world large-scale applications. A comprehensive analysis has been conducted, implementing an audio and video model, and observed that this solution has outperformed most traditional methods of violence detection on the XD-Violence Dataset, as evidenced by the accuracy metrics.

Read the paper · More papers on PaperTik