Spoken word recognition of independent speakers using local‐peak‐weighted average dictionary
Hiroshi Matsumoto, Masao Nakagawa, Masahide Yoneyama · Electronics and Communications in Japan (Part I Communications) · 1986
Abstract A method has been proposed in the spoken word recognition, which utilizes the frequency‐time pattern inherent to the word. However, in the application of the method for speech by unspecified speakers, normalization is required, in contrast to the case of specified speakers, since the formant in the pattern may be shifted along the frequency‐axis due to the individual difference of harmonization organs of the speakers. The method including the formant normalization or nonlinear conversion of the frequency‐axis contains a larger computational complexity, and is not suited to a large vocabulary of words. From such a viewpoint, this paper proposes a method, in which a common weighted average dictionary is constructed by superposing the binary local‐peak patterns of speakers, absorbing their individual differences into a dictionary. The individual differences of speakers are also overcome by matching with the dictionary the binary pattern with an allowable width along the frequency‐axis centered around the local peak of the input speech. The weighted average dictionary is only a pattern set representing the frequency of the local peaks inherent in the word, and the storage of data requires less memory capacity. Since the input pattern is matched based on the binary pattern set, the computational complexity is also reduced. With 110 words of 0A command as the sample speech, both male and female speech is used as the inputs. When the input speakers are restricted, a recognition rate of 97 percent was obtained for restricted speakers and nearly 92 percent was obtained for unrestricted speakers.