Speech modeling and voiced/unvoiced/mixed/silence speech segmentation with fractionally Gaussian noise based models

Shayan Oveisgharan, Mohammad Bagher Shamsollahi · 2004

The ARMA filtered fractionally differenced Gaussian noise (FdGn) model and a new AR filtered FdGn added up model are applied to a speech signal and performance of their parameters on speech unvoiced/voiced/mixed/silence classification is evaluated against the zero crossing rate (ZCR) feature. For parameter estimation of AR filtered FdGn model two methods were applied: the iterative maximum likelihood (ML) method of Tewfik (1993) and a new computationally efficient linear minimum square error (LMSE) algorithm. Also for parameter estimation of the new added up model two approaches were implemented: an expectation-maximization (EM) based approach and an iterative MSE approach. The described models and methods were applied to a speech signal and also its real cepstrum. The performance of the described models on V/U/M/S speech classification was obtained based on the J/sub 1/ parameter in this order: added up model on real cepstrum of speech, filtered FdGn model on real cepstrum of speech (LMSE method), filtered FdGn model on speech (LMSE method), ZCR, and filtered FdGn model on speech (Tewfik method).

Read the paper · More papers on PaperTik