Survey of current speech technology
Alexander I. Rudnicky, Alexander G. Hauptmann, Kai-Fu Lee · Communications of the ACM · 1994
This article describes two technologies, speech recognition and speech synthesis, that manipulate speech in terms of its information content. Recognition is the transformation of human speech into text to be used literally (e.g., for dictation) or interpreted as commands to control applications. Synthesis allows the generation of spoken utterances from text. Synthesis is desirable when a large number of utterances must be available or when message content is unpredictable, requirements that make pre-recording of speech impractical. The technologies covered in this article are of particular interest because they support direct communication between humans and computers through a mode that humans commonly use for communication amongst themselves and at which they are highly skilled. Other speech technologies of note, not discussed here, include speaker recognition (automatically establishing a speaker's identity) as well as speech editing and indexing (the manipulation of speech without the extraction of linguistic information). 1 RECOGNITION