Pretrained network-based sound event recognition for audio surveillance applications
Suwan Park, Geonwoo Kim · 2021 International Conference on Information and Communication Technology Convergence (ICTC) · 2021
Despite the recent surge in the demand on the audio recognition in surveillance systems, there are still many obstacles to readily use it in real environments such as legal restrictions on public data collection and difficulties in obtaining large-scale learning data. To overcome these problems, we propose an adaptive sound event recognition scheme based on a pre-trained network with the large-scale AudioSet, where PANNs-based CNN and SincNet[3] are employed to extract the audio features from log-mel spectrogram and waveform, respectively. Our experimental results show that the proposed method achieves mean average precision (mAP) of 0.415, which is slightly better than the best previous methods. Furthermore, we evaluate the performance of transfer leaning using a smaller amount of data collected by ourselves, considering dangerous situation scenarios in surveillance applications.