A Synthesized Voice Discrimination Method Using Characteristic Sounds Spoken by Humans
Yuya Nuruki, Yoshiaki Taniguchi · 2023
In this paper, to prevent victimization of crimes such as impersonation fraud using synthesized voice, we propose a method to determine whether the speaker is a human or a synthesized voice using characteristic sounds spoken by humans. In our proposed method, the input voice data is first segmented into regular intervals, and the segmented data is further converted into spectrogram images. The converted images are classified into the four characteristic sounds defined in this paper using a CNN. Using the relationship between breath sounds, saliva sounds, and speech duration in human voice and synthesized voice, which were investigated in advance, the speaker of the input voice data is discriminated as human or synthesized voice. We evaluate our proposed method using actual human voice and synthesized voice generated by Japanese speech synthesis tools, and show that the proposed method shows high discrimination results.