Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
Heng Cui, Longfei Song, Li Li, Dongxing Xu, Yanhua Long · Journal of King Saud University - Computer and Information Sciences · 2025
Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal model selection and integration for SED. We propose a framework combining heterogeneous SSL model (BEATs, HuBERT, WavLM, etc.) representations through three fusion strategies: individual SSL embedding integration, dual-modal fusion, and full aggregation. Our experiments on the DCASE 2023 Task 4 Challenge reveal that dual-modal fusion (e.g., CRNN+BEATs+WavLM) achieves complementary performance gains, while CRNN+BEATs alone attains superior individual SSL embedding integration results. We further introduce a normalized sound event bounding boxes (nSEBBs), an adaptive post-processing method that dynamically adjusts target events detection boundaries, improving PSDS $$_1$$ by up to 4% for standalone SSL models. These findings provide important insights into SSL model compatibility, demonstrating that task-specific fusion and dynamic post-processing enhance robustness. Our work establishes a reference for selecting and integrating SSL architectures in SED systems, balancing efficiency and accuracy for real-world deployment.