Automatic Recognition of Spoken Numerals
George Sebestyen · The Journal of the Acoustical Society of America · 1960
A spoken numeral is represented by a vector in a 361 dimensional space constructed through use of an 18-channel Vocoder. In the quantized time-frequency plane representation of speech events each dimension expresses the amount of energy the spoken utterance contains in a given band of frequencies during a given interval of time. Utterances are reduced to a fixed duration, and the time normalization factor is included as one of the dimensions of the space. Linear transformations of the space (one for each of the class of digits zero through nine) are found by a computer so that, after transformation, a given set of vectors (points) which represent different utterances of the same digit become maximally clustered in the new space. Membership in the ten classes of spoken numerals is measured by ten different non-Euclidean metrics obtained from the Euclidean metric via the linear transformations. Thus different error criteria, each generated by members of a class of numerals, are used to measure the membership of an utterance in the ten classes of numerals. This technique was tried on 400 samples of the ten numerals spoken by ten speakers. Results of the experiment indicate that no errors in recognition were made when seven or more samples were used to compute the optimum linear transformations. This result was obtained when the numerals tested were spoken by persons other than those whose samples were given originally.