A Packet-Based CAPDM Speech Coder for
Chia-Horng Liu, Chia-Chi Huang · 2000
In this paper, we present a median-rate speech coder, the controlled adaptive prediction delta modulation coder (CAPDM), which operates at 16 kb/s with good speech quality and low algorithm complexity (15). The coder is dedicated to personal communication network (PCN) applications and transmits speech samples on the basis of packets. It combines the features of one-step looking forward decision, syllabic companding, instan- taneous companding, and adaptive prediction. In addition to the use of a short-term prediction filter, CAPDM also exploits the pitch property to predict speech waveform explicitly. With the aid of a pitch prediction filter, the performance of a CAPDM codec improves about 3 dB in segmental signal-to-noise ratio (SEGSNR). The average SEGSNR of CAPDM.FF is about 21 dB, which is 7 dB over traditional CVSD at 16 kb/s. We also utilize an adaptive postfilter (APF) to enhance the perceptual quality of the decoded speech. The mean opinion score (MOS) listening test of CAPDM.FF with APF shows that its average score achieves 4.19, which is as good as G.728 16-kb/s LD-CELP and is comparable with CCITT G.721 32-kb/s ADPCM. The complexity of CAPDM.FF is evaluated to be 8 MIPS, which is much lower than that of LD-CELP and could be further reduced by adopting a smaller correlation window for pitch detection. To solve the problem of packet loss, we developed a packet-based waveform substitution method by reinitializing the codec parame- ters at the beginning of each packet. The simulation results show that CAPDM.FF could tolerate 5% of packet loss and still keep an SEGSNR at 10 dB and an MOS at about 3.0. activity detection (VAD). In this paper, we propose a 16-kb/s speech codec with appropriate speech quality for PCN applica- tions. The adoption of this speech codec will double the PCN system capacity as compared with using the G.721 ADPCM codec. Furthermore, if VAD mechanism is employed with this codec, much higher system capacity gain can be achieved. The structure of a controlled adaptive prediction delta mod- ulation coder (CAPDM) codec consists mainly of four parts: a logic unit, a stepsize estimation unit, a pitch predictor, and an adaptive prediction unit. The logic unit is used for one-step look forward decision. The stepsize estimation unit is used to esti- mate both instantaneous and syllabic stepsizes. The pitch pre- dictor is used to find out vowel pitch periods. The adaptive pre- diction unit is used to predict the current speech sample based on both a short- and a long-term predictor. For PCN applications, digital transmission schemes are used to carry voice or data packets over a shared radio medium (26). When a network is in a heavy traffic, a voice packet might be held in a queue. When a voice packet is delayed over a certain time limit, it must be dropped. Another problem of packet voice transmission comes from packet loss. In order to sustain speech codec performance in an error prompt environment, people use different waveform substitution techniques to recover lost voice packets. In this scenario, we have surveyed the zero substitu- tion technique (3), the previous packet repetition technique (3), the pattern matching technique (3), the pitch-based replication technique (4), etc. (5). Besides waveform substitution, the conti- nuity of codec parameters must be maintained in order to bridge over the gaps of the lost packets. In the paper, we show that the isolation of each transmission packet will avoid the problem of divergence in the reconstructed speech waveform and maintain good codec performance when voice packets are lost. This paper is composed of six parts. In Section II, we de- scribe the basic structure of a CAPDM codec. In Section III, we describe two pitch prediction methods which can be used in a CAPDM coder. In Section IV, we present computer sim- ulation results of CAPDM codec in an ideal channel in which both subjective and objective performance are evaluated. In Sec- tion V, we describe CAPDM performance evaluation in a noisy channel. Finally, we draw some conclusions in Section VI.