Multi-Modal Speech Separation Based on Two-Stage Feature Fusion

Yang Liu, Ying Wei · 2021 IEEE 6th International Conference on Signal and Image Processing (ICSIP) · 2021

How to efficiently achieve speech separation has always been a difficult problem for computers. In this paper, we propose a multi-modal speech separation based on two-stage feature fusion to solve the cocktail party problem. The model makes full use of the audio information and the speakers’ visual information as input, and a two-stage feature fusion strategy is proposed. Inspired by the characteristic of the cochlea, a multi-stream convolution neural network is used to extract high-frequency and low-frequency audio features, respectively. Then time convolution network is applied to extract fused audio features formed by the connection of high-frequency features and low-frequency features in the first stage of fusion. The extracted audio feature is connected with the visual feature from the dilated convolutional layers, which is the second stage feature fusion. To fully exploit the amplitude and phase information of the audio, we choose the time-frequency mask in the complex domain as the training target of the model. Through a series of experiments, we prove that the performance of our model is superior to other methods. Besides, our model is more suitable for application in actual scenes due to its low complexity.

Read the paper · More papers on PaperTik