Speaker Identification Using Cepstrum in Kannada Language

C. Srividya, S. Savithri, Physics Section · 2011

Speaker identification in forensic cases involves careful estimation of the features that are more specific to speaker. The anatomical differences in vocal tract depict speaker related differences. Cepstrum has been investigated as a possible parameter for speaker identification. Fundamental frequency obtained from cepstral coefficients indirectly depicts the shape and size of vocal tract. In the present study quefrency and amplitude was extracted using cepstrum spectral analysis technique under various recording conditions. 30 normal males were selected. For reading intention, four paragraphs were selected having long vowels /a:/ , /i:/ and /u:/ embedded in the words of the paragraphs of Kannada passage. To elicit spontaneous speech from the subjects, six Kannada words having long vowels in the medial position were considered. Subject’s speech and reading paragraphs were recorded in field conditions to suit realistic forensic situations. Using CSL-4500 software, cepstral coefficients quefrency and amplitude were extracted for the long vowels /a:/, /i:/ and /u:/. Extracted parameters were normalized and Euclidian distances between speakers were measured. The results indicated that the percent correct identification was above chance level for direct vs. direct and mobile vs. mobile recording (DS Vs. DS = 68%, MS Vs. MS = 64%, DR Vs. DR = 75%, and MR Vs. MR = 62%). Results of 4-way repeated measures ANOVA revealed that significant difference between speakers on quefrency and interaction between speaking style, recording, set and vowels on quefrency. Results of paired t-test reveals the significant difference between direct and mobile recording for long vowels /a:/, /i:/ and /u:/ on quefrency.

Read the paper · More papers on PaperTik