Auditory/visual speech in multimodal human interfaces
Dominic W. Massaro, Michael M. Cohen · 1994
It has long been a hope, expectation, and prediction that speech would be the primary medium of communication between humans and machines. To date, this dream has not been realized. We predict that exploiting the multimodal nature of spoken language will facilitate the use of this medium. We begin our paper with a general framework for the analysis of speech recognition by humans and a theoretical model. We then present a system for auditory/visual speech synthesis that performs complete text-to-speech synthesis. This system should improve the quality as well as the attractiveness of speech as one of a machine's primary output communication medium. Mirroring the value of multimodal speech synthesis, multimodal channels should also enhance speech recognition by machine. 1. INTRODUCTION Speech perception is a human skill that rivals our other impressive achievements. Even after decades of intense effort, speech recognition by machine remains far inferior to human performance. Our thesis ...