Bone Conduction-Aided Speech Enhancement With Two-Tower Network and Contrastive Learning

Changtao Li, Feiran Yang, Jun Jie Yang · IEEE Transactions on Audio Speech and Language Processing · 2024

This paper presents a time-domain two-tower network, namely BiNet, that jointly utilizes bone-conducted speech and noisy air-conducted speech for multi-modal speech enhancement. The presented BiNet adopts two independent encoders to map bone-conducted speech and noisy air-conducted speech into a shared embedding space. Subsequently, the decoder in BiNet reconstructs the target clean speech using the embedding features from both modalities. Compared to the widely adopted single-encoder networks, the presented two-tower structure can fully exploit the information intrinsic to each modality. To effectively fuse local and global speech features, we incorporate skip connections between the two encoders and the decoder in BiNet. We utilize a multi-scale mel-spectrogram loss function originally proposed for speech synthesis as the training objective for BiNet. Moreover, the two-tower structure of BiNet prompts us to leverage contrastive learning-based regularization. By controlling the similarity between the embedding features of bone-conducted speech and noisy air-conducted speech, we consider two types of regularization constraints on the two encoders of BiNet, and find that BiNet performs better as the embedding features of both modalities exhibit a higher similarity. Extensive experiments demonstrate that the proposed method outperforms single-modal and multi-modal speech enhancement systems significantly in terms of perceptual evaluation of speech quality (PESQ) and short-time objective intelligibility (STOI).

Read the paper · More papers on PaperTik