Deciphering speech waveforms

M. O'Kane, J. Gillis, Phil Rose, M. Wagner · 2005

Many phoneticians are remarkably expert at 'reading' speech waveforms. This paper describes an attempt to capture this knowledge for use as a segmentation and early labelling knowledge source for a continuous speech recognition system. As well as deriving information from the waveform directly, the decisions made by the waveform deciphering knowledge source are based on a related series of functions derived from the waveform. These functions, which relate to both valley-to-peak and zero crossing measures, are computationally very efficient and it would seem that the frequency analogues of these functions could provide an alternative means of deriving a certain amount of the spectral information more usually obtained through spectrograms.

Read the paper · More papers on PaperTik