Temporal decomposition and acoustic-phonetic decoding for the automatic recognition of continuous speech

Paul Deléglise, Frédéric Bimbot, Claude Montacié, Gérard Chollet · 2003

Localization is realized using a robust implementation of temporal decomposition, a technique originally proposed for speech coding. Speech is decomposed in terms of overlapping events characterized by spectral targets and time-limited interpolation functions. The spectral targets are evaluated iteratively. An undershot target may be reestimated using neighbors and the associated functions. The possibility of undoing the effects of coarticulation is also suggested. The identification of these corrected targets is therefore possible with no further contextual rules. Applications to speech recognition strategies are discussed. The recognition of spelled surnames (letters of the alphabet) is used for evaluation. The correct recognition of 76% of correct phones is the basis of 70% correct recognition letters.>

Read the paper · More papers on PaperTik