Cues in the perception of infinitely clipped speech

Daniel Kahn · The Journal of the Acoustical Society of America · 1985

“Infinitely clipped” (“IC”) speech is surprisingly intelligible [Licklider and Pollack, J. Acoust. Soc. Am. 20, 42–51 (1948)]. Because infinite clipping so severely distorts the speech spectrum, it is sometimes speculated that a time-domain similarity between natural and IC speech—precise identity of zero-crossing locations—might account for the latter's intelligibility. This hypothesis has practical implications (should speech recognition algorithms look beyond power spectra?) and theoretical ones [cf. Scott, J. Acoust. Soc. Am. 60, 1354–1365 (1976)]. However, the hypothesis would be refuted if it turned out that the small amount of residual similarity between the power spectra of the original and IC signals could alone account for the observed intelligibility. Three types of stimuli—natural vowels, /p,t,k/ in natural-sentence context, and synthetic vowels—were infinitely clipped (“IC stimuli”), then subjected to transformations which preserved the power spectrum but grossly distorted the phase spectrum, thereby destroying the original zero-crossing information. These transformed IC stimuli were no less intelligible than the unaltered IC stimuli, strongly suggesting that power spectrum rather than zero-crossing cues are responsible for IC-speech intelligibility. Further, constant-parameter, synthetic-vowel inteiligibility decreased dramatically upon clipping, natural-vowel intelligibility relatively little, demonstrating the usefulness of spectral change and/or prosodic cues in the absence of accurate formant information.

Read the paper · More papers on PaperTik