Neuronal oscillations in decoding time-compressed speech
Oded Ghitza · The Journal of the Acoustical Society of America · 2016
At the core of oscillation-based models of speech perception is the notion that decoding is guided by parsing. In these models, parsing is executed by setting a time-varying, hierarchical window structure synchronized to the input. Syllabic parsing is into speech fragments that are multi-phone in duration, and it is realized by a theta oscillator capable of tracking the input syllabic rhythm, with the theta cycles aligned with intervocalic speech fragments termed theta-syllables. Prosodic parsing is into fragments that are multi-word in duration, and it is realized by a delta oscillator capable of tracking phrase-level prosodic information, with the delta cycles aligned with chunks. Intelligibility remains high as long as the oscillators are in sync with the input, and it sharply deteriorates once they are out of sync. In the pre-lexical layers, decoding is realized by a cascade of neuronal oscillators in the theta, beta, and gamma frequency bands, with theta as “master.” This talk reviews a model that utilizes this cortical computation principle, capable of explaining counterintuitive data on the intelligibility of time-compressed speech hard to explain with conventional models of speech perception. [Work supported by AFOSR.]