Modelling Spatio-Temporal Dynamics by Graph Attention Network for Distributed Multi-Microphone Sound Event Classification

Vijay John, Yasutomo Kawanishi · 2025

This paper introduces a novel distributed multi-microphone sound event classification framework that uses graph attention networks to model spatial and temporal relationships between distributed multi-microphones. Existing methods utilize naive aggregation approaches like concatenation, averaging, or maximum operations for multiple sensor inputs resulting in information loss and limitations in capturing the complex spatial and temporal relationships in the multi-microphone sequence. To address this, we propose a framework based on the graph attention network to model the complex multi-microphone relationship. Our framework is based on two graph structures: a spatio-temporal graph (STG), which is algorithmically modeled to capture inter-microphone and inter-frame relationships, and a learnable fully-connected spatial graph (FCSG), which is designed to capture complementary details. Utilizing them, a graph attention network-based aggregation module effectively updates the graph nodes resulting in improved event classification accuracy. Experimental results on the MM-Office dataset demonstrate that our proposed framework significantly outperforms baseline methods for the event classification task.

Read the paper · More papers on PaperTik