Speech Enhancement for Punjabi Language Using Deep Neural Network

Jaspreet Singh, Kamaldeep Kaur · 2019

Speech contains sounds that delineate words of a language. Capturing speech in different real environmental circumstances may contain different mixtures of many different sound waves (i.e. the mixture of sounds, noises and other effects like reverberation). Such noisy environments make difficult to capture clear speech and hence degrade the speech quality. Due to the different nature of noise and reverberation effect, it must be necessary to remove extra noise and reverberation effect separately. This paper proposes a hybrid approach that uses various algorithms sequentially to remove noise, reverberation effect, and cocktail party problem. Non-negative matrix factorization (NMF) is representation learning method that uses basis matrix and activation coefficients to produce the basic spectral structure of speech. Deep learning use supervised mapping function to map noisy spectrum to target speech spectrum. NMF with deep learning can jointly be used to separate non-stationary noise from target speech. On other hands, the non-negative convolutive transfer function (nCTF) estimates the coefficient of reverberation. Non-negative convolutive transfer function with NMF can be used to reduce the effect of the room impulse response (RIR). Gaussian mixture model-universal background model (GMM-UBM) extracts the compact representation of speech. GMM-UBM with deep classifier used to extract target speech from cocktail party problem (the mixture of sounds of multiple speakers). The Proposed system is based on the Punjabi language that aims to enhance speech under noisy-reverberant environments and the cocktail party problem.

Read the paper · More papers on PaperTik