X-STA: Cross-Modal Spatial-Temporal Alignment Network for Unified Audio-Visual Segmentation
Hanyu Xuan, Tongxing Liu, Wenxiang Dong, Zhongheng Li, Shuo Chen · IEEE Signal Processing Letters · 2025
Audio-Visual Segmentation (AVS) aims to segment sound sources from video frames using synchronized audio cues. This task requires not only localizing the sound sources within frames but also accurately delineating their shapes. Existing AVS methods often rely on assumptions of spatial-temporal consistency between audio-visual content and are typically designed for specific learning paradigms. However, this specialization limits their ability to handle multi-granularity supervision signals and and adapt to diverse task requirements. For this purpose, we propose a Cross-modal Spatial-Temporal Alignment (X-STA) network to alleviate spatial-temporal inconsistency and overcome paradigm-specific constraint. Our X-STA introduces three key components: a novel multi-stage Cross-modal Adapter (xAdapter) that transfers knowledge from a pre-trained SAM through multi-grained representation adaptation, an innovative Cross-modal Prompter (xPrompter) that provides geometry-aware constraints for AVS through dynamic prompting strategies, and a Cross-modal Self-supervised (xSelf) mechanism that refines temporal alignment and enables self-supervised AVS. These components collectively facilitate explicit reasoning about location and geometric shape of the sound source by refining the alignment of cross-modal spatial-temporal cues. Our method achieves competitive performance across several baselines on widely-used AVS datasets, demonstrating its effectiveness in addressing the complexities of AVS.