Improved Methods for Pitch Synchronous Linear Prediction Analysis of Speech
麗清 劉 · Institutional Repositories DataBase (IRDB) · 2015
Linear prediction (LP) analysis has been applied to speech system over the last few decades. LP technique is well-suited for speech analysis due to its ability to model speech production process approximately. Hence LP analysis has been widely used for speech enhancement, low-bit-rate speech coding in cellular telephony, speech recognition, characteristic parameter extraction (vocal tract resonances frequencies, fundamental frequency called pitch) and so on. However, the performance of the conventional LP method is degraded by high-pitched harmonic structure of glottal excitation source and background noise. In order to improve the performance of LP analysis, it is necessary to reduce the effect of these two factors, which is a most challenging task for LP analysis. The objective of this dissertation is to develop some approaches to improve the performance of the LP analysis based on pitch synchronous analysis. We consider a pitch synchronous LP analysis for high-pitched speech using a weighted short time energy (STE) function for the purpose of downgrading the effect of the harmonic structure of the glottal excitation source. Unlike some conventional techniques, which require the electroglottography (EGG) signal or complicated epoch extraction algorithms, we utilize a simple STE computation of speech signal and prediction residual signal to extract the interval of glottal closed phase during a glottal cycle and do not need to estimate the instant of glottal closure and opening exactly. To reduce the infuence of the background noise, we propose a noise compensation LP method based on pitch synchronous analysis under white noise environment. Exploiting the periodicity of voiced speech and random distribution of background white noise, a more accurate estimation of noise power is calculated on each current frame of speech. The advantage, that the noise power is estimated from each current frame, can avoid the estimation delay and accuracy problem. Sometimes the background noise could be white or colored signals. A noise whitening method for the noise compensation LP Method is proposed so that the new noise estimator can be also applied to colored environment. We further propose a crosscorrelation sequence-based LP analysis under noisy environment. The crosscorrelation sequence is utilized to replace the original speech signal which is sensitive to background noise, and applied to LP analysis. The approach can improve the performance of LP analysis under noisy environment. In this dissertation, we focus on resolving the two factors that degrade the performance of LP analysis and new approaches have been proposed and implemented. The experimental results, based on synthetic and real speeches, demonstrate the effectiveness of the new approaches for improving the performance of the LP analysis.