Speaker normalization using cortical strip maps: A neural model for steady state vowel identification
Heather Ames, Stephen Grossberg · The Journal of the Acoustical Society of America · 2007
Auditory signals of speech are speaker-dependent, but representations of language meaning are much more speaker-independent. Such a speaker normalization transformation enables speech to be learned and understood from different speakers. A neural model is presented that performs speaker normalization to generate a pitch-independent representation of speech sounds, while also preserving information about speaker identity. This speaker-invariant representation is categorized into unitized speech items, which input to sequential working memories whose distributed patterns can be categorized, or chunked, into syllable and word representations. The proposed model circuits fit into an emerging theory of auditory streaming and speech categorization in which auditory streaming and speaker normalization both use similar neural designs; namely, multiple cortical strip map representations of auditory signals. This design homology may clarify how speaker normalization circuits evolved from more primitive streaming mechanisms. Simulations with synthesized steady-state vowels from the Peterson and Barney (1952) vowel database achieve accuracy rates similar to those achieved by human listeners. These results are compared to behavioral data and other speaker normalization models. [Work supported in part by NSF and ONR.]