Speech Enhancement and Encoding using SS-VAD and LPC
Thimmaraja Yadava G, H.C. Vinay, T.R. Nayana, H. S. Jayanna, P. Lavanya, D. Aswini, Garima Singh G. · 2019
In this paper, an enhancement of noisy speech data and encoding of corrupted and enhanced speech data is presented. An algorithm is proposed which is an amalgamation of spectral subtraction with voice activity detection (SS-VAD) and linear predictive coding (LPC) for speech enhancement and encoding purpose. Firstly, the significance of SS-VAD and LPC methods are studied in detail for various types of noises. In SS-V AD technique, the noisy speech data is considered as an input speech signal which is a combination of clean speech data and noise model. The corrupted speech data is windowed using Hanning window and framed for every 20ms. 50% of overlapping is done while windowing the speech data. The output of SS-VAD is given as an input to the LPC encoder. The coefficients are extracted from the input speech data to design all pole filters. The cross correlation process is also done for differentiating the voiced and unvoiced samples at the analysis step. The pitch information and extracted coefficients are used at the synthesis step. The experiments are conducted for different types of noisy speech data which are degraded by background noise, F16 noise, factory noise, white noise and car noise. The experimental results show that an SS-VAD algorithm significantly improved the signal to noise ratio (SNR) of speech data under various degraded conditions. Therefore, the SS-VAD algorithm is combined with LPC for encoding of enhanced speech data which yields better audibility and intelligibility of speech compared to encoding of noisy speech data.