A Video Vision Transformer for Sound Source Localization

Haruo Yokota, Mert Bozkurtlar, Benjamin Yen, Katsutoshi Itoyama, Kenji Nishida, Kazuhiro Nakadai · 2024

This paper addresses sound source localization (SSL) based on a data-driven approach using neural networks. Data-driven approaches can achieve superior performance by training models on data recorded in the target environment, compared to traditional signal processing approaches. Two problems make it difficult to achieve high-performance SSL with a data-driven approach: 1) models can not handle properly the temporal context to balance robustness of SSL and dynamic changes like moving sources, and 2) are not highly expressive and accurate enough to learn SSL tasks to handle periodic phase information, which is the key feature to SSL. To solve these problems, we propose a stride division method of the acoustic features to handle the temporal context appropriately. In addition, we propose to introduce a video vision transformer (ViViT), which is known to achieve high-performance video recognition, to SSL. Experimental results using real acoustic signals showed that the proposed method outperforms conventional methods, demonstrating the effectiveness of the proposed method.

Read the paper · More papers on PaperTik