Spoken Language Identication With Hierarchical Temporal Memories
Dan R. Robinson, Kevin Leung, Xavier Falco · 2009
The task of spoken language identification has enjoyed many promising approaches via machine learning and signal processing since the late 1970s. However, most either employ Hidden Markov Models (HMMs) to model sequential data or use language-dependent phoneme recognizers as primary features [3, p. 33]. The former has sparsity-related issues and is sensitive to previously unseen events, which are common in human language. The latter requires extensive labeling of training data on a phonological or prosodic level, which is often impractical, and depends upon specialized information for each spoken language under consideration. Our project is an attempt to use Hierarchical Temporal Memory (HTM) to do spoken language identification. HTM, implemented and distributed by the Numenta company, is a new technology inspired by the human neocortex. Distinguishing characteristics of HTMs include their ability to natively process both temporal and spatial information and their potential for deep hierarchical structure. These two features of HTMs suggest their utility in audio and language tasks.