FASR: Effect of voice disguise

Kesiya Sebastian, Leena Mary · 2016

Speaker recognition systems based on spectral features perform well in acoustically matched and noise-free conditions. Spectral features are unsuccessful to model information about the speaker at higher levels. Prosody which represents intonation, rhythm and stress of speech, better represents speaker characteristics at higher levels. Voice disguise is a common threat to automatic speaker identification systems. It is the action of transforming one's voice purposefully to imitate or just to conceal one's identity. This work attempts to study the effects of voice disguise on prosody based speaker verification system. For this a prosody based automatic speaker verification system is implemented using support vector machine (SVM). For the extraction of prosodic features, speech is segmented into phrase-like regions by detecting long pauses regions and further syllable-like segmentation is done at the valleys of short time energy contour. In order to segment fused syllables, vowel onset point (VOP) detection is performed in such regions. Then prosodic features are extracted and used for building SVM models. To study the effect of voice disguise, normal and masked speech were recorded simultaneously for a set of speakers in the lab environment. Prosodic features for both normal and masked speech were analyzed. Using normal speech, automatic speaker verification system is implemented and testing was done using both normal and masked speech. Through this study, it is identified that prosodic features derived from pitch and duration are almost robust against voice masking but energy features are slightly affected.

Read the paper · More papers on PaperTik