Speaker Information using Subsegmental and Segmental Analysis of LP Residual
Debadatta Pati · 2009
Abstract—Linear Prediction (LP) residual mostly contains the excitation source information. This work analyzes the LP residual once using frame size of 5 ms (subsegmental) and another time using frame size of 20 ms (segmental), each with a shift of 2.5 ms. The residual frames are then subjected to non-parametric Vector Quantization (VQ) to store the unique excita-tion sequences for each speaker. The testing of such codebooks seem to contain significant speaker information. Further the speaker information at the two levels are found to be different in nature. The performance of the speaker recognition system using subsegmental, segmental and combined subsegmental-segmental speaker information is found to be 71.67%, 55 % and 83.33%, respectively, for a population of 30 speakers taken from TIMIT database using a codebook of size 128. Further combination of the proposed subsegmental-segmental LP residual based system with the vocal tract system feature based system providing stand alone performance of 95 % is found to be 100%. This aspect reinforces the different speaker information available in the excitation component of speech. Index Terms: speaker specific information, excitation source, LP residual, subsegmental, segmental