Schematizing spectrograms for speech recognition

Michael Riley · The Journal of the Acoustical Society of America · 1983

An initial step in wide vocabulary, continuous speech recognition is proposed that roughly consists of a schematization of the features seen in conventional spectrograms—e.g., peaks and edges of spectral energy concentrations, temporal discontinuities, and spectral balance information. In a second step, these features can be mapped onto their acoustic-phonetic correlates—e.g., formant distribution, voice onsets, articulatory closures. A method for locating spectral energy concentrations is given that takes advantage of their usual continuity in time, and thus performs superior to locating peaks in spectral cross sections. It begins with smoothing and flattening convolutions in both the time and frequency dimensions of narrow-band spectrograms to select the appropriate temporal and spectral scales. Ridges in the resulting two-dimensional (time-frequency) surfaces correspond to local spectral energy concentrations. The tops of these ridges are found by the application of a two-dimensional differential operator at each point in the time-frequency plane. The operator's definition in terms of the relationship between the gradient and principal directions will be given, along with justification and examples.

Read the paper · More papers on PaperTik