Speech Signal Processing

Jonathan Stein · 2000

In this chapter we treat of one of the most intricate and fascinating signals ever to be studied, human speech. Here our knowledge of these mechanisms to the practical problem of speech modeling is applied. Speech synthesis is the artificial generation of understandable and natural-sounding speech. If coupled with a set of rules for reading text, rules that in some languages are simple but in others quite complex, we get text-to-speech conversion. We introduce the reader to speech modeling by means of a naive, but functional, speech synthesis system. Speech recognition, also called speech-to-text conversion, seems at first to be a pattern recognition problem, but closer examination proves understanding speech to be much more complex due to time warping effects. Although a difficult task, the allure of a machine that converses with humans via natural speech is so great that much research has been and is still being devoted to this subject. There are also many other applications—speaker verification, emotional content extraction (voice polygraph), blind voice separation (cocktail party effect), speech enhancement, and language identification, to name just a few. While the list of applications is endless many of the basic principles tend to be the same. We focus on the deriving of ‘features’, i.e., sets of parameters that are believed to contain the information needed for the various tasks. Simplistic sampling and digitizing of speech requires a high information rate (in bits per second), meaning wide bandwidth and large storage requirements. More sophisticated methods have been developed that require a significantly lower information rate but introduce a tolerable amount of distortion to the original signal. These methods are called speech coding or speech compression techniques, and the main focus of this chapter is to follow the historical development of telephone-grade speech compression techniques that successively halved bit rates from 64 to below 8 Kb/s.

Read the paper · More papers on PaperTik