Timbre features for speaker identification of whispering speech: selection of optimal audio descriptors

Vijay M. Sardar, Suresh Damodar Shirbahadurkar · International Journal of Computers and Applications · 2019

Whispered mode of speech is tailored for sharing the confidential information or to maintain the peace at public place like a library. Sometimes it is used designedly by the criminals to hide the identity. The speaker identification from the whispered voice has gained lot importance in forensic science, as a whisper is tough to mimic. The phonation which owes the prominent intelligence to discriminate the speakers is missing in a whispered speech to its neutral counterpart. Hence, it is difficult to spot the person from its whisper voice by machine using the traditional features. Human identifies the person from its whisper voice by the perceptual sense, hence a perceptually motivated timbre features are used here. Only well-performing timbre features namely Roll-off, Brightness, Irregularity, Roughness and, MFCC (Mel Frequency Cepstral Coefficient) are selected for the task using Hybrid Selection Algorithm. An enhancement of about 5.8% is reported in the identification accuracy by the proposed system compared to a baseline using the state of the art NDMP (non-parametric density modeling) based fusion-SVM system.

Read the paper · More papers on PaperTik