Forensic synthetic speech inspection technique based on formant comparison
Qimeng Lu, Puyang Geng, Hong Mei Guo, Xiaohong Chen, Shaopei Shi · 2022
With the development of speech synthesis technology, the simulation of specific individual’s speech has gradually matured, synthetic speech is easily and perceptually recognized as real speech, which may occur frequently in illegal activities. To identify crimes, forensic technology is widely used such as comparing the formants, pitches, and rhythm. The present study aims to investigate whether the method of comparison of formants can recognize the differences between the perceptually similar synthetic speech (hereinafter “personal anchor” speech) and real speech. To this end, two young males and two young females from different dialect regions were recruited to read the same text. Their voices were recorded and used to generate four “personal anchors” by the software of sound spectrum and statistics analysis. The method of comparison of various parameters of the formant, including numerical statistical, stability analysis, and transitional segments feature were applied to analyze the differences between the real speech and the corresponding “personal anchors”. It was found that the numerical or stability analysis of formants was not sufficient to fully determine whether the speech was synthesized, while comparing the transitional segments of some specific syllables could efficiently detect the synthetic speech from the real speech.