A 9600 bit per second speech compression algorithm based on linear prediction
B. Abzug · 1981
The use of digital technologies in voice communication systems has created the need for efficient methods of converting analog speech signals into digital data formats. Often this conversion must be accomplished subject to a fixed bit rate constraint. For rates below 16,000 bits per second (bps), the speech data is usually processed by a compression algorithm to preserve the quality and intelligibility of the speech. This dissertation presents a new 9600 bps speech compression algorithm. The new algorithm is based on a state-of-the-art 2400 bps linear predictive coder (LPC). A synthesizer excitation is derived from the prediction residual and is used to replace the fixed impulse excitation used in LPC. A method for selecting, evaluating, and coding the excitation is presented. A simulation of the new algorithm was developed to evaluate the synthesized speech. Quality and intelligibility tests were performed. They demonstrated that the use of the residual-derived excitation increases the naturalness and intelligibility of the synthesized speech relative to the baseline. The performance of the new algorithm was also compared to that of a 9600 bps adaptive predictive coder (APC). For male speakers, the speech synthesized by the new algorithm was generally preferred, while for female speakers, APC was judged superior. The new algorithm contains features that makes it particularly suited for applications where both 2400 bps LPC and a 9600 bps compression algorithm was required. For example, the bit stream of the new algorithm is a super set of the LPC bit stream. This greatly facilitates conferencing between users who have only the low rate capability and users whose compression systems operate at the higher rate. Also, much of the software required to implement the new algorithm is identical to that of LPC. A 9600 bps capability can be added to a 2400 bps LPC compression system with only a modest increase in program memory.