Automatically derived units for segment vocoders
Viswanathan Ramasubramanian, T.V. Sreenivas · 2004
Segment vocoders play a special role in very low bitrate speech coding to achieve intelligible speech at bitrates of /spl sim/300 bits/sec. We explore the definition and use of automatically derived units for segment quantization in segment vocoders. We consider three automatic segmentation techniques, namely, spectral transition measures (STM), maximum-likelihood (ML) segmentation (unconstrained) and duration-constrained ML segmentation, towards defining diphone-like and phone-like units. We show that the ML segmentations realize phone-like units which are significantly better than those obtained by STM in terms of match accuracy with TIMIT phone segmentation as well as actual vocoder performance measured in terms of segmental SNR. Moreover, the phone-like units of ML segmentations also outperform the diphone-like units obtained using STM in early vocoders. We also show that the segment vocoder can operate at very high intelligibility when used in a single-speaker mode.