Enhanced Automated Audio Captioning Method Based on Heterogeneous Feature Fusion

Liwen Tan, Yi Zhou, Yin Liu, Wang Chen · 2024

Automated Audio Captioning (AAC) serves as a bridge between audio and text modalities by generating descriptive captions for audio data. However, the inherent challenge posed by modality differences continues to hinder research progress in this field. Recent studies focus on uncovering intrinsic connections between modalities by learning their similarities and differences. However, this approach often overlooks useful information across different samples. This paper proposes an AAC method based on heterogeneous feature fusion, which introduces the combined Local-Global feature Fusion module (LGFuser). By fusing heterogeneous audio features, the method captures high-dimensional semantic information, effectively bridging the gap between audio and text modalities. On the AudioCaps and Clotho V2 datasets, the method achieves METEOR, CIDEr, and SPIDEr-FL scores of 0.187, 0.475, and 0.303, respectively.

Read the paper · More papers on PaperTik