Disambiguating speech commands using physical context
Katherine M. Everitt, Susumu Harada, Jeff Bilmes, James A. Landay · 2007
Speech has great potential as an input mechanism for ubiquitous computing. However, the current requirements necessary for accurate speech recognition, such as a quiet environment and a well-positioned and high-quality microphone, are unreasonable to expect in a realistic setting. In a physical environment, there is often contextual information which can be sensed and used to augment the speech signal. We investigated improving speech recognition rates for an electronic personal trainer using knowledge about what equipment was in use as context. We performed an experiment with participants speaking in an instrumented apartment environment and compared the recognition rates of a larger grammar with those of a smaller grammar that is determined by the context. Figure 1: We tagged gym objects with modified RFID tags to provide context to the speech recognizer. Categories and Subject Descriptors the device. Speech itself is commonly used and so requires little H.5.2 [User Interfaces]: Voice I/O additional training to use. However, current speech recognizers are often inaccurate in non-controlled conditions due to ambient General Terms