Speech coding a new approach
Shyamal Kr. Das Mandal, A.K. Datta, S. Chowdhury · 2004
Text-to-speech synthesis, based on ESNOLA, uses signal dictionary having raw sound signals representing parts of phonemes. State-phase analysis for detection of voiced region along with detection of pitch also may be used for extraction of the most appropriate signal elements automatically from continuous speech in real time. The signal elements at the voiced zone are perceptual-pitch-periods. These signal are coded by simply inserting one information byte at the beginning of each element. The decoding is done using the information bit. The intervening signals are regenerated by linear estimation from the two perceptual-pitch-periods. This coding induces a ten-fold information reduction without significant loss of naturalness.