Incorporating linguistic knowledge and automatic baseform generation in acoustic subword unit based speech recognition

Trym Holter, Torbjørn Karl Svendsen · 1997

A major challenge in speech recognition based on acoustic subword units is creating a lexicon which is robust to inter- and intra-speaker variations. In this paper we present two different approaches for incorporating simple word-level linguistic knowledge into the labelling step of the training procedure. The proposed systems also utilise a scheme for combined optimisation of baseforms and subword models. For the TI46 database, these methods are shown to greatly improve the performance compared to an acoustic subword based speech recogniser employing unsupervised labelling, and they are found to perform as well as systems utilising whole-word models and context independent phoneme models. 1. INTRODUCTION Traditionally, automatic speech recognisers employ phone-like units based upon a linguistic description of the language. On the other hand, the analysis of the actual speech signal is acoustically based. The resulting system is neither phonetically nor acoustically consistent, but is...

Read the paper · More papers on PaperTik