Advanced time-frequency displays applied to forensic speaker identification.
Sean A. Fulop, Sandra Ferrari Disner · The Journal of the Acoustical Society of America · 2009
Forensic phoneticians have long utilized conventional spectrograms as the basis of the the auditory-acoustic method of speaker identification. Recent investigations [S. A. Fulop and K. Fitz, “Using the reassigned spectrogram to obtain a voiceprint,” J. Acoust. Soc. Am. 119, 3337 (2006); S. A. Fulop and S. Disner, “The reassigned spectrogram as a tool for voice identification,” Int. Congress of Phonetic Sciences XVI, 1853–1856 (2007)] have suggested that a newer technology, i.e., the reassigned spectrogram, may be at least as valuable in the service of speaker identification. The present paper describes methods for forensic speaker identification using enhanced time-frequency displays including reassigned spectrograms and Zhao-Atlas-Marks distributions. Identification methods which employ the aforementioned displays are compared with the more established auditory-acoustic method, using recordings of 10 speakers saying several short phrases two times each. The results point to the efficacy of these newer time-frequency displays for speaker identification, particularly when used to corroborate conventional methods. One important development is the ability to focus on the detailed time-frequency features of the phonation process, thereby obtaining a new speech biometric indicative of vocal cord dynamics.