A reflective infrastructure for building spoken-language interfaces

Michael S. Fulkerson, Alan W. Biermann · 2001

This thesis presents JAVOX a novel infrastructure that allows traditional desktop applications to be transformed into spoken-language systems (SLSs). Historically, weaknesses in speech recognition and natural-language processing (NLP) have been blamed for the dearth of applications with speech interfaces; however these technologies have progressed to a point where they can support very usable, limited-vocabulary applications. Until now, the majority of SLSs have been built as a side effect of research into a particular aspect of spoken interfaces; very little attention has been paid to SLSs themselves. Ad hoc designs, myriad architectures, and a lack of formal approaches have led to the popular opinion that building SLSs is inherently hard. Further, because most of the historical SLSs have used designs that vary greatly from similar systems without a voice interface, the general assumption has been that SLSs require fundamentally different designs than their non-speech-enabled counterparts. In this research, I have assumed the opposite: speech-enabled programs should use the same designs as would similar applications without speech, and adding a spoken-language interface should be no harder than building the underlying application. The JAVOX infrastructure presents a principled and formal approach for building desktop SLSs. It allows programs that were written without a speech capability to be speech enabled with, in most cases, no source-code modifications. JAVOX uses the reflective capabilities found in many object-oriented languages to see into and control executing programs. The speech and language processing capabilities of JAVOX operate beneath and are transparent to the target application. The instructions for translating from spoken language to executable code are contained in a J AVOX grammar file. The JAVOX grammars are made up of two distinct, yet complementary, languages: the JAVOX Scripting Language and Semantically Extended Java Speech. The JAVOX infrastructure has a general mechanism for the processing of multimodal interactions. It allows developers to support pointing behaviors without the need for timestamps or complicated coordination schemes. The general infrastructure, coupled with generic multimodal capabilities, makes JAVOX an ideal test platform for NLP and human-computer interaction (HCI) research. Researchers can quickly begin investigating their primary topic, without spending time developing speech systems from the ground up. The power of the JAVOX infrastructure is demonstrated through the use of numerous examples, several of which were developed in under two days—a small fraction of the usual time required to speech enable an application.

Read the paper · More papers on PaperTik