The difference between acoustic and auditory parameter signals as a cue for phonetic segmentation and categorization
Hans G. Tillmann, Bernd Pompino-Marschall · The Journal of the Acoustical Society of America · 1988
Normally, speech recognition systems are based on purely acoustic or on auditorily modeled acoustic speech signals, respectively. The difficulties in segmentation and categorization encountered by those models may partly be overcome by a technique that exploits the difference signal between acoustic and auditory parameters. The difference signal tested here was between the sound-pressure envelope and the loundness time series. Sound-pressure level (dB-scaled) was calculated as the mean value of absolute amplitude within a rectangular widow of 300 sample points (= 15 ms). Loudness was computed every 15 ms according to Paulus and Zwicker [Acustica 27, 253–266 (1972)] by using critical band amplitudes approximated by averaged DFT values (with frequency-dependent differences in length of the Hanning window; 1500 Hz: 15 ms). For calculation of the difference, signal loudness and sound-pressure level were normalized to 0 < × < 1 and subtracted from one another.