TFT: An algorithm for the spectral compression of natural speech signals

Richard R. Hurtig · The Journal of the Acoustical Society of America · 1989

An earlier report [R. R. Hurtig, J. Acoust. Soc. Am. Suppl. 1 81, S78 (1986)] using synthesized syllables demonstrated that naive subjects can discriminate and identify spectrally compressed vowel segments under auditory and vibrotactile conditions. These findings are consistent with the view that the identification of the spectral shape of the speech segment may be independent of its frequency range. A computational algorithm was developed to achieve spectral compression of natural speech. The algorithm includes calculation of an n-point FFT, padding the result with the spectrum of a Hamming window, calculation of a 2n-point IFFT, and outputting the first half of the resultant time domain signal. The size of the pad determines the amount of compression achieved while the placement of the pad determines the direction of the frequency shift. Naive subjects had no difficulty recognizing simple sentences in a closed set for speech signals compressed to 2500 or 1250 Hz bands. After a few hours of listening, open set recognition is achieved. The implementation of the algorithm for sensory aids will be demonstrated.

Read the paper · More papers on PaperTik