Affine Transformations of the Speech Space
M. Pijpers, Mike Alder · 1992
The papers Speaker Normalization of static and dynamic vowel spectral features (J.A.S.A 90, July 1991 pp 67-75) and Minimum Mean-Square Error Transformations of Categorical Data to Target Positions (IEEE Trans Sig.Proc,40 Jan 1992, pp13-23) by Zahorian and Jagharghi describe an algorithm for transforming the space of speech sounds so as to improve the accuracy of classification. Classification was accomplished by both back-propagation neural nets and by a Bayesian Maximum Likelihood method on the model of each vowel class being specified by a gaussian distribution. The transformation was an affine transformation obtained by choosing ideal `target' points for each cluster in a second space and minimising the mean square distance of the points in the speech space from the appropriate target. The speech space itself was a space of cepstral coefficients obtained from a Discrete Cosine Transform. These findings are remarkable, indeed almost unbelievable. The reason is that both the maximum...