Recognition of spoofed voice using convolutional neural networks

Huixin Liang, Xiaodan Lin, Qiong Zhang, Xiangui Kang · 2017

As people use spoofed voice commonly, the dilemma of speech recognition (SR) occurs frequently, while it is critical to determine whether a voice is disguised. Nevertheless, the detection for less distorted voices, especially for those with the disguising factors of ±4 semitones that sound like a natural person and that are more popularly employed in the methods of disguise, is not satisfactory. In order to solve this problem, we propose an approach using the convolutional neural network (CNN) to identify the spoofed voice by pretreatment analysis of the electronically disguised speech. Experimental results show that CNN can effectively solve the declining detection in the pitch of ±4 semitones, in which the accuracy is higher than 95% with an advantage gap of 4.2% over the best conventional method. Besides, the cross-database identification results also perform better than the traditional methods, indicating that CNN has a significant potential on SR.

Read the paper · More papers on PaperTik