Vowel formant tracking for speaker identification
Ming Feng Jiang, Min Shi · The Journal of the Acoustical Society of America · 1991
The first two or three spectral peaks, or formants, are crucial in determining the vowel quality. In turn, accurate determination vowel formants quality is important to effective speaker identification task. A vowel formant tracking vector (VFT) was developed for the speaker identification (SAUSI) profile. Specifically, the speech spectrum is obtained frame-by-frame by using an LPC algorithm with the first three formant frequencies for each frame calculated. The underlying assumption was that the vowels will exhibit a contiguous formant frequency transition from frame-to-frame and, hence, can be separated from consonants for the cited formant measurements. In order to carry out this task, the frequency range 0–5000 Hz is divided into 34 semitone bins and three histograms are obtained for first three vowel formants. In turn, these histograms provide an estimation of general quality of the vowels spoken by each speaker being evaluated. The result is that the interspeaker differences are large enough to permit identification of the target speaker while the intraspeaker differences are fairly small even for text independent speech. The algorithm utilized will be presented as will data demonstrating that this VFT vector is robust enough to effectively perform the speaker identification task.