Comparing Classification Methods in Isolated Vowel Classification
Ondrej Šuch, Santiago Barreda, Martin Klimo, Peter Tarábek, Andrea Tinajová · 2018
Frame-wise processing of speech is the standard approach in automated speech recognition. The problem of assigning correct posterior classification probabilities to individual frames is quite subtle, because there is no ground truth available. It can be approached only indirectly, by considering a classification or recognition task requiring classification of a longer sequence of frames and inferring quality of posteriors from its success. In our paper we consider the problem of classification of isolated vowels in the TIMIT corpus with the ultimate goal of judging the quality of frame posteriors. We design a novel two level model for this purpose. At the first, spectral level, the posterior distribution is computed via multiclass classification based on the spectrum of each analysis frame. At the second, temporal level, temporal analysis of the first level posteriors is performed and serves as the input to a second level classifier. As an application of this model we compare multiple pairwise coupling classifiers. Experimental results indicate that currently preferred coupling methods based on the Bradley-Terry model give poorer quality posteriors than a simpler coupling method based on averaging.