Some practical considerations in the deployment of a wireless-communication interactive voice response system

Carmén García Mateo, Laura Docío-Fernández, Antonio Cardenal-López · 2001

Abstract In this paper, we describe the design procedure for a wirelesscommunication interactive voice response system. The applica-tion must work in a very noisy environment which has imposedmany design constraints. Wewilladdress the sensible aspects ofthree components of the application: the voice activity detector(VAD),the automatic speech recognition (ASR)system, and theconfidence measure (CM) determination. In order to get a sat-isfactory product, it has been necessary to reduce the importantmismatch between available linguistic and acoustic resourcesand the operational environment. Adaptation techniques for theacoustic models of the speech recognition system have provento be effective to speed up the application deployment time. 1. Introduction It is widely accepted that speech technology is nowadays ma-ture enough for the deployment of practical, commercial ap-plications in the real-world. In the last years, many servicesand products have shown up specially in the telecommunicationarena[1]. Also, many resources as speech and text databases,recognition engines, text-to-speech engines are available to thedevelopers for building up applications and products [2]. Nev-ertheless, the concourse of highly specialized and well-trainedpersonnel is required in order to get a quality product. Even af-ter a careful design of the project specifications and an adequateselection of linguistic and software components, an error-and-trial strategy must be conducted in order to tune and debug allthe components. This is even more delicate when there is a sig-nificantmismatch between theavailable linguisticresources andthe actual acoustic and linguistic environment. This is the caseof the project we will describe in this paper. The applicationitself, a command-control system, is simple in nature, but theacoustic environment (a very noisy lumberyard), the transmis-sion channel (a half-duplex radio channel), and the two possiblelanguages (Galician and Spanish) are factors that make very dif-ficult the transfer of technology from off-the-shelf componentsto a complete product.A state-of-the-art system may perform poorly when the testdata are collected under a totally different environmental con-dition. Regarding to the possible mismatches, both linguisticand acoustic mismatches might occur. A linguistic mismatchis mainly caused by incomplete task specification, inadequateknowledge representations, and insufficient training data. Onthe other hand, an acoustic mismatch between training and op-erational conditions arises, being in this case changes in trans-ducers, channel, speaker environment, background noise, etc.We will address how some techniques must be applied inorder to improve the performance and ergonomy of the appli-cation. We have selected three subsystems that required of spe-cial attention: the voice activity detector (VAD), the automaticspeech recognition (ASR) system, and the confidence measure(CM) determination. Before describing the most relevant as-pect of each of them, we will show an overview of the wholeapplication. Some conclusions will be drawn at the end of thepaper.

Read the paper · More papers on PaperTik