Direct identification vs. correlated models to process acoustic and articulatory informations in automatic speech recognition

Régine Andre-Obrecht, Bruno Jacob · 2002

Our work deals with the classical problem of merging heterogenous and asynchronous parameters. It is well known that lip reading improves the speech recognition score, especially in noise conditions; so we study more precisely the modeling of the acoustic and labial parameters to propose two automatic speech recognition systems: a direct identification is performed by using a classical HMM approach, no correlation between visual and acoustic parameters is assumed; and two correlated models, a master HMM and a slave HMM, process respectively the labial observations and the acoustic ones. To assess each approach, we use a segmental pre-processing method. Our task is the recognition of spelled French letters, in clear and noisy (cocktail party) environments. Whatever the approach and conditions, the introduction of labial features improves the performance, but the difference between the two models is not enough sufficient to provide any priority.

Read the paper · More papers on PaperTik