Revealing the processing history of pitch-shifted voice using CNNs
Lihua Wang, Huixin Liang, Xiaodan Lin, Xiangui Kang · 2018
With the emergence of varieties of audio editing software, disguised voices can be generated easily and are able to spoof the automatic speaker verification system. In this work, we focus on the spoofed voice disguised by pitch shifting which is commonly used by popular audio editing software and design a forensic chain to reveal the processing history of pitch-shifted voice signals blindly based on convolutional neural networks (CNNs). Because tracing the disguising tool before recognizing the disguising factor from the spoofed speech helps much to recognize the disguising factor, the proposed framework enables the detection of the disguised voice and the utilized tools. To our knowledge, it is the first attempt, focusing on voice transformation without any prior knowledge of the original speech, to recognize the disguising tool and the disguising factor from the spoofed speech, which is more suitable in real-world applications. Experimental results show that the disguising factors can be recognized with accuracy higher than 90%, indicating a high possibility of restoring the original voice from its disguised counterpart.