Speech recognition by indexing and sequencing

Simone Franzini, Jezekiel Ben-Arie · 2010

Recognition by Indexing and Sequencing (RISq) is a general-purpose method for classification of temporal vector sequences. We developed an advanced version of RISq and applied it to isolated-word speech recognition, a task most commonly performed with Hidden Markov Models (HMMs) or Dynamic Time Warping (DTW). RISq is substantially different from both these methods and presents several advantages over them: robust recognition can be achieved with only a few samples from the input sequence and training can be carried out with one or more examples per class. This enables much faster training and also allows to recognize speech with a variety of accents. A two-step classification algorithm is used: first the training samples closest to each input sample are identified and weighted with a parallel algorithm (indexing). Then a maximum weighted bipartite graph matching is found between the input sequence and a training sequence, respecting an additional temporal constraint (sequencing). We discuss the application of RISq to speech recognition and compare its architecture and performance with that of Sphinx, a state-of-the-art speech recognizer based on HMMs.

Read the paper · More papers on PaperTik