Machine Learning Approaches for Audio Classification in Video Surveillance: A Comparative Analysis of ANN vs. CNN vs. LSTM
Shriganeshrajkumar S. Togare, Ashwini Andurkar · 2023
This study aims to develop a robust audio classification system capable of accurately analyzing and categorizing audio events in video streams. Leveraging machine learning techniques, we construct a model that can recognize and classify various audio events commonly encountered in surveillance scenarios, including gunshots, sirens, and car horns.Audio samples are extracted from video frames, and crucial audio features are enhanced while removing noise during the preprocessing stage. Feature extraction techniques, such as Mel-frequency cepstral coefficients (MFCC), are employed to capture distinctive audio traits. A labeled dataset comprising 8,732 audio samples is utilized to train a machine learning model, such as a Convolutional Neural Network (CNN) or Artificial Neural Network (ANN) OR long short-term memory networks(LSTM). The model progressively learns to identify and categorize each audio pattern and characteristic. To improve the model's performance and generalizability, diverse methods are applied. Real-world video surveillance datasets are employed in multiple experiments to assess the proposed method's performance. The developed audio classification system is compared with existing approaches to demonstrate its superiority. The results reveal that the implementation of the audio classification system significantly enhances surveillance capabilities by accurately identifying and categorizing audio events in video streams. This capacity improves the overall effectiveness and efficiency of surveillance systems, enabling operators to promptly recognize and respond to potentially harmful situations.