PROSODY-BASED TEXT INDEPENDENT SPEAKER IDENTIFICATION METHOD
Irena Chmielewska · 2004
In the paper/poster, a speaker identification method exploiting the envelope of the speech signal fundamental frequency waveform (prosody) as a resource of the individual distinctive speaker-related characteristics is presented. The method is based on spectral analysis of the test utterances developed by recording two different texts read consecutively by a set of ten speakers (male and female voices). The method involves extraction of the speech signal fundamental frequency (Fo) from the recorded speech, determination of the Fo intensity waveform within the utterances of given duration time and statistical analysis of changes in the Fo intensity waveform, the changes being treated as the variable. In the Fo extraction process, the segmentation of utterances is applied. After having calculated the statistical parameters for the Fo value distribution, two-dimensional speaker’s identification model in the form of the Speaker Identification Area (SIA) is developed for each speaker belonging to the test group. Such a model forms ‘a speechprint ’ of the individual speaker. Regarding the form of speech signals being the would-be inputs in the real speaker identification process, the method was tested with the band-limited (70- 1400 Hz) speech signal input. As the identification error rates are comparable, the individual speechprints seem to be located in the lower band of the speech signal spectrum.