Multiple Frame Rates for Feature Extraction and Reliable Frame Selection at the Decision for Speaker Identification Under Voice Disguise

Swati Prasad, Zheng‐Hua Tan, Ramjee Prasad · Journal of CONASENSE · 2016

Determining the person who spoke a given speech utterance from a group of people is referred to as Speaker Identification.It is used in crime scenes, surveillance and consumer electronic products like smart TV.But it faces poor performance due to a mismatch between the train and the test speech data, that arises because of the adoption of voice disguise.Therefore, this paper studies the effect of three different types of voice disguises, namely, Fast (nonimitative), Synchronous (Imitative) and Repetitive Synchronous Imitation along with the normal speaking from the CHAINS corpus on the speaker identification performance.Finally, a system combining different frame rates for feature extraction and reliable frame selection at the decision level has been proposed.The evaluated system showed an overall better performance than the baseline systems.

Read the paper · More papers on PaperTik