TFFNN:Time-Frequency Fusion Neural Network for Voice Antispoofing
Jiapan Gan, Xiaodan Lin · 2024
In the voice antispoofing task, it is crucial to identify and analyze the key factors that affect the detection performance. The purpose of this paper is to explore the relative importance of time-domain features versus frequency-domain features in the antispoofing task. Firstly, we compare the effectiveness of these two types of features in recognizing spoof speech based on a one-dimensional time-delayed neural network (TDNN) and a residual network (ResNet). Secondly, combined time domain and frequency domain features, the Time-Frequency Fusion Neural Network, i.e., TFFNN, is proposed for voice antispoofing, the architecture firstly uses a non-standard two-dimensional convolution kernel aiming at focusing on the time-frequency domain features at the same time, but with the frequency domain features as the main focus, and then specializes in one-dimensional convolution processing for the time domain, and finally fuses them in the channel dimensions, through this unique approach, the TFFNN is able to effectively extract time-frequency features from speech signals, thus enhancing the performance of anti-fraud systems. The experimental results show that although the frequency domain features show significant advantages in capturing the subtle changes of speech, the time domain features that capture the overall rhythm and dynamic range of speech cannot be neglected. We find that the combined effect of time-domain and frequency-domain features exhibits significant advantages in enhancing the performance of antispoofing task on the ASVspoof2019 dataset.