A novel algorithm for sparse classification.
Sujeeth Bharadwaj, Mark Hasegawa‐Johnson · The Journal of the Acoustical Society of America · 2010
A recent result in compressed sensing (CS) makes it possible to perform non-parametric speech recognition that is robust to noise and that requires few training examples. By taking fixed length representations of training examples and stacking them in a matrix, a frame or an over-complete basis can be constructed. Gemmeke and Cranen showed that sparse projections onto this frame recover the correct transcription with 91% accuracy at -5-dB SNR. The goal of speech recognition is not sparse projection onto training tokens, but onto training types. Sparse projection onto types can be achieved by building a frame for each word in the dictionary and stacking the frames to form a rank 3 tensor. Speech recognition is performed by convex linear projection onto the tensor, with sparsity enforced only in the index that specifies type. A mixed L1/L2 relaxation was derived, leading to a Newton descent algorithm that converges to the global optimum.