A Sequential Audio Spectrogram Transformer for Real-Time Sound Event Detection

Takezo Ohta, Yoshiaki Bando, Keisuke Imoto, Masaki Onishi · 2024

In this paper, we propose an audio spectrogram transformer (AST) for sequential inference and evaluate its real- time performance. ASTs are pre-trained in a self-supervised manner, such as masked autoencoding, and the pre-trained models are well-performing in sound event detection. However, the existing architectures are designed for offline inference, wherein the entire signal serves as the input, and are unsuitable for sequential inference as they require the input sequence to be split into short chunks. In this study, we design a sequential AST based on a memory token (MT-AST) and its training method and conduct comprehensive experiments regarding the chunk length configuration. Specifically, we extend the offline AST with special tokens that memorize past signal information so that the network avoids repetitive inference of the same signal. While our model has limited inference capability, we train it using knowledge distillation from BEATs, a large-scale pre-trained model. Compared to the offline architecture, our model achieved higher performance by pre-training with AudioSet and fine-tuning for the URBAN-SED and DESED datasets. In addition, we conducted experiments to investigate the input chunk length considering performance-latency trade-offs and revealed the optimal configurations. We revealed that our model requires at least one extra second of input to maintain the performance.

Read the paper · More papers on PaperTik