Text-Independent Speaker Recognition using Subsegmental, Segmenetal & Suprasegmental Features

Kuldeep Kumar, G. Venkateswarlu, Tingting Rao, G. Chenchamma · 2011

Current speaker recognizer systems the speaker specific source information at different levels. In this we exploits the source information (LP residual) present at different levels namely subsegmental, segmental & suprasegmental. The subsegmental analysis considers LP residual in blocks of 5 msec with shift of 2.5 msec to extract speaker information. The segmental analysis extracts speaker information by processing in blocks of 20 msec with shift of 2.5 msec. The suprasegmental speaker information is extracted by viewing in blocks of 250 msec with shift of 6.25 msec. The speaker recognizer studies performed using TIMIT (Texas Instruments and Massachusetts Institute of Technology) databases demonstrate that the segmental analysis provides best performance followed by subsegmental analysis. The suprasegmental analysis gives the least performance. However, the evidences from all the three levels of processing seem to be different and combine well to provide improved performance, demonstrating different speaker information captured at each level of processing. Finally, the combined evidence from all the three levels of processing together with vocal tract information further improves the speaker recognition performance. I. Introduction The Speaker information can be classified by two types Speaker Identification (SI) and Speaker Verification (SV). SI & SV based on text spoken in two types Text Dependent and Text Independent. Speaker Identification is One-to-Many as well as The Speaker Verification is One-to-One. The speaker information in the speech signal is based on the physiological behavioral of a person(Atal 1972). The physiological aspects are due to the vocal band and excitation source that involved in the production. The behavioral aspect involves aspects like speaking rate, accent, voice frequency etc.The shape, size and the dynamics are related with the vocal band contribute to the speaker characteristics. In similar way, the shape, size and the dynamics associated with the vocal folds contribute to the speaker characteristics. State of the art speaker recognition systems mostly use vocal tract related speaker information represented by the spectral or cepstral features like linear prediction cepstral coefficients (LPCC) or mel frequency cepstral coeffi cients (MFCC) These features provide good recognition performance. The reason may be that, they nearly represent complete vocal tract information. By that we mean, LPCC or MFCC captures the formants and their bandwidth information characterizing the vocal tract completely, but pitch is only one aspect of speaker information due to source. The 5 msec blocks of LP residual sample sequences in the time domain are used as feature vectors for modeling speaker information by Gaussian mixture modeling (GMM) technique to generate subsegmental speaker models. The 20 msec blocks of LP residual samples are first decimated by a factor of 4 to reduce its dimensionality and also to eliminate the information that has been modeled at the subsegmental level. The decimated LP residual sample sequences are modeled by GMM to generate segmental speaker models. The decimated LP residual sample sequences are modeled by GMM to generate suprasegmental speaker models. All these models are independently tested using respective blocks of LP residual extracted from the test signals to evaluate the amount of speaker information present at each level. Finally the combination of evidences from all the three levels is made to observe their different nature of speaker information. The potential of combined source information is demonstrated by comparing and also combining its performance with a speaker recognition system using vocal tract feature. If we treat the residual signal as random noise, then the distribution of the samples will be Gaussian. Since the LP residual deviates from random noise due to pitch information, to that extent the distribution of the residual samples may be non-Gaussian in nature. However, this can be handled with the help of the GMM. Hence the motivation for using GMM for speaker modeling from LP residual. The rest of the paper is organized as follows: de-scribes the proposed subsegmental, segmental and suprasegmental analysis of LP residual approach for modeling speaker information from the LP residual. In this paper will also describe the speaker recognition studies that have been performed using the proposed approach. describes an alternative approach for suprasegmental information using the concept of instantaneous pitch. The last section summarizes the present work with a mention on the scope for future work.

Read the paper · More papers on PaperTik