Biologically Plausible Speech Recognition Using S pike-Based Phase Locking Cues (Invited Paper - Special Session)
Ismail Uysal, John G. Harris · 2009
A biologically plausible algorithm is proposed for phoneme recognition that makes use of spikes for computation. The prototype system is demonstrated on voiced phonemes and shows competitive performance with state-of-the-art systems on a vowel dataset in the presence of noise. Using novel phase locking cues, the algorithm performs surprisingly well down to SNR values as low as 5dB, where the performance of the baseline algorithm used for comparison drops considerably. These results suggest, not only a possible explanation for the extreme noise robustness of human recognition, but also a novel technique for information processing for speech signals with computers. I. INTRODUCTION Since human speech recognition capabilities significantly outperform state-of-the-art automatic speech recognition (ASR) systems, it is natural that many researchers have in- cluded biological inspiration in the design of ASR algorithms. In fact, the most popular ASR feature set, the Mel frequency cepstral coefficients (MFCC), makes use of the psychoacoustic properties of the auditory system, such as the non-uniform distribution of the auditory filters throughout the frequency range of hearing (1). However, these filters are merely a pre- processing step in the complex hierarchical structure of the auditory system where most of the computation is carried out via action potentials that originate from the cochlea as spike trains traveling through the auditory nerve (AN) fiber. There has been some limited research in spike train rep- resentations for speech recognition (2), (3), however, these algorithms either use novel, yet biologically non-plausible ways to generate spike trains from speech or use the spike trains that are generated in the AN fiber connected to the cochlea without any spike encoding schemes. We believe that better understanding of the information encoding of AN fiber spike trains will lead to more noise-robust ASR systems. This research explores various spike encoding schemes as possible feature sets to come up with a noise-robust, spike-based feature domain for speech recognition. Furthermore, a fully spike- based classification algorithm is developed to employ the new feature set for a comparison with state-of-the-art ASR on a noisy vowel dataset. Our results show that phase locking cues inherent in the AN spike trains could be used as a competitive feature set for ASR and might provide a possible explanation for the superb robustness of the auditory system. Pattern recognition problems usually consist of two fun- damental steps: the extraction of a feature set which forms a compact and robust representation of the input signal; and classification in the generated feature space. Section II will describe the standard and spike-based methods for feature extraction for speech. Rank order coding and liquid state machine as spike-based classifiers will be discussed in Sec- tion III. Section IV includes the performance tests and results. Section V concludes the paper.