Voicing detection in DAP-STC
M.S. Ho, D.J. Molyneux, Barry M.G. Cheetham · 2002
Sinusoidal transform coding (STC) requires an all-pole representation of spectra derived periodically from the short-term speech spectral envelope and a "voicing probability" frequency f/sub v/ to divide each spectrum into two sub-bands: voiced below f/sub v/ and unvoiced above f/sub v/. Discrete all-pole (DAP) modeling may be applied to STC to improve the accuracy of the short-term spectral envelope for voiced speech with modifications to accommodate unvoiced speech and spectra which do not conform well to an all-pole model. This paper presents a novel approach to the determination of f/sub v/ which is appropriate when DAP is employed. It is a frequency-domain algorithm with an analysis-by-synthesis optimisation process. This approach improves the accuracy of DAP-STC modeled speech.