Multi-microphone periodicity function for robust F0 estimation in real noisy and reverberant environments
Federico Flego, Maurizio Omologo · 2006
Abstract This paper outlines a new method to extract F0 from distant-talking speech signals acquired by a microphone network, whichexploits the redundancy across the signals proceeding from eachmicrophone, by jointly processing the different contributes. Tothis purpose, a multi-microphone periodicity function is derivedfrom the magnitude spectrum computed on each microphone sig-nal. This function allows to estimate F0 reliably, even under re-verberant conditions, without the need of any post-processing orsmoothing technique. Experiments, conducted on real lectures,showed that the proposed frequency-domain algorithm is moresuitable than other time-domain based ones. IndexTerms : speechanalysis,fundamentalfrequencyestimation,multi-microphone processing, distant-talking interaction. 1. Introduction In the CHIL project, various signal processing techniques are be-ing investigated that aim to address challenging problems amongwhich acoustic event classification, speakerlocalization and track-ing, distant-talking speech recognition, speech activity detection,speaker identification and verification [1].One way to pursue all these objectives is that of deriving amodel of the source (e.g. the speaker) from the given multi-microphone data. To this purpose, a Distributed Microphone Net-work (DMN) is used, which consists in a generic set of micro-phones localized in space without any specific geometry.In this work we address the problem of deriving a robust es-timation of the fundamental frequency F0 from the variety of sig-nals recorded through the microphone network. Speech signalsrecorded by microphones placed far from a talker are severely de-graded by both background noise and reverberation, which de-pends on spatial relationships among the microphones and thetalker, as well as on the scenario acoustic characteristics.Estimating F0 independently for each microphone signal andapplying then majority vote or other fusion based methods mayrepresent a possible approach. Another way to perform F0 es-timation is to extend to the multi-microphone case a paradigmthat works for a single microphone close-talking case. A time-domain F0 extraction algorithm based on Weighted Autocorre-lation (WAUTOC) [2] was experimented in the past [3], whichshowed good performance on a real multi-microphone databaseof distant-talking speech sequences reproduced in an office envi-ronment. In particular, the resulting multi-microphone WAUTOCtechnique offers the advantage of obtaining better performance