Spatial speech coding for multi-teleconferencing

Kok Soon Phua, Woon‐Seng Gan · 2003

This paper describes a structural model for the implementation of multichannel speech coding for teleconferencing with spatial audio reproduction. Multiple nonaural speech sources are synthesized into binaural sound to produce a more realistic videoconferencing environment. The activity information of the individual binaural speech, which is determined by the voice activity detection algorithm, is used to calculate two weighting factors prior to mixing. Furthermore, a third level of weight adjustment can be carried out by adjusting these weighting factors before applying to the individual voice source. A scheme to remove undesirable noise spikes is also introduced. Both channels are then coded using the G.723.1 speech codec individually.

Read the paper · More papers on PaperTik