Phonetic Classification Using Hierarchical, Feed-forward, Spectro-temporal Patch-based Architectures

Ryan M. Rifkin, Jake Bouvrie, Ken Schutte, Sharat Chikkerur, Minjoon Kouh, Tony Ezzat, Tomaso Poggio · 2007

... object recognition was applied to the task of phonetic classification. During learning, the system processed 2-D wideband magnitude spectrograms directly as images, producing a set of 2-D spectrotemporal patch dictionaries at different spectro-temporal positions, orientations, scales, and of varying complexity. During testing, features were computed by comparing the stored patches with patches from novel spectrograms. Classification was performed using a regularized least squares classifier (Rifkin, Yeo et al. 2003; Rifkin, Schutte et al. 2007) trained on the features computed by the system. On a 20-class TIMIT vowel classification task, the model features achieved a best result of 58.74 % error, compared to 48.57 % error using state-of-the-art MFCC-based features trained using the same classifier. This suggests that hierarchical, feed-forward, spectro-temporal patch-based architectures may be useful for phonetic analysis.

Read the paper · More papers on PaperTik