Noise robust signal processing for human pitch tracking and bird song classification and detection

Abeer A. Alwan, Wei Chu · 2012

This dissertation investigates the extraction of discriminative information from noisy signals for human fundamental frequency (FO) tracking, and for bird song classification and detection. For FO tracking, the investigation is carried out in the direction of reducing FO estimation and voicing decision errors. To reduce FO estimation errors, a novel Statistical Algorithm for FO Estimation, SAFE, is proposed to improve the accuracy of FO estimation under both clean and noisy conditions. Prominent Signal-to-Noise Ratio (SNR) peaks in speech spectra constitute a robust information source from which FO can be inferred. A probabilistic framework is proposed to model the effect of noise on voiced speech spectra. Prominent SNR peaks in the low frequency band (0 – 1000 Hz) are important to FO estimation, and prominent SNR peaks in the middle and high frequency bands (1000 – 3000 Hz) are also useful supplemental information to FO estimation under noisy conditions, especially in the babble noise condition. To reduce voicing decision errors, we introduce a model-based unvoiced/voiced (U/V) classification frontend which can 1w used by any FO tracking algorithm. We propose an FO Frame Error (FFE) metric which combines Gross Pitch Error (GPE) and Voicing Decision Error (VDE) to objectively evaluate the performance of FO tracking methods. A GPE-VDE curve is then developed to show the tradeoff between GPE and VDE. For bird call classification, the investigation is carried out in the direction of signal denoising and discriminative feature extraction. To enhance noisy signals, we propose a Correlation-Maximization denoising filter which utilizes periodicity information to remove additive noise in Antbird calls. We also develop a statistically-based noise-robust bird-call classification system which uses the denoising filter as a frontend. Enhanced bird calls which are the output of the denoising filter are used for feature extraction. To obtain discriminative features, we extend the expectation-maximization (EM) algorithm to estimate not only optimal acoustic model parameters, but also optimal center frequencies and bandwidths of the filter bank used in cepstral feature extraction for bird call classification. The search is done using the gradient ascent method. Filter tank and model parameters are optimized iteratively. For bird song detection, temporal, spectral, and structural characteristics of Robin songs and syllables are studied. Syllables in Robin songs are clustered by comparing a distance treasure defined as the average of aligned Linear Predictive Coding (LPC)-based frame level differences. The syllable patterns inferred from the clustering results are used to improve acoustic modelling of a hidden Markov model (1- MM)-based song detector.

Read the paper · More papers on PaperTik