Automatic robust classification of speech using analytical feature techniques
G. Pérez · 2009
This document reports the research done in the domain of automatic classification of speech within a Master’s degree internship in the Sony CSL laboratory. The work explores the potential of the EDS system, developed at Sony CSL, to solve speech recognition problems of a small number of isolated words, independently of the speaker, and with the presence of background noise. EDS automatically builds features for audio classification problems. This is done by means of (functional) composition of mathematical and signal processing operators. These features are called analytical features and are built by the system specifically for each audio classification problem, given under the form of a train and a test database. In order to adapt EDS to speech classification, since features are generated through functional composition of basic operators, a research on specific operators for speech classification problems has been done, and new operators have been implemented and added to EDS. To test the performance of our approach to the problem, a speech database has been created, and experiments before and after adding the new specific operators have been carried out. An SVM classifier using EDS analytical features has then been compared to a standard HMM-based speech recognizer. The results of the experiments indicate, on the one hand, that the new operators have shown to be useful to improve the speech classification performance. On the other hand, they show that EDS performs correctly in a speaker-dependent context, while further experimentation has to be done to draw conclusions in a speaker-independent situation.