Speaker localization based on oriented global coherence field

Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer · 2006

Abstract This paper proposes a new speaker localization method that isbased on a preliminary estimation of the head orientation. The ba-sic information on which the estimation is accomplished is calledOriented Global Coherence Field (OGCF).The new algorithm is shown to be significantly more robustthan the traditional ones so far explored. Its robustness is also dueto an effective speech activity detection, implicitly performed bya thresholding technique applied to OGCF information. To showthe performance of the proposed system, experiments were con-ducted on the NIST RT-05 Spring Evaluation source localizationtask, which is based on real recordings of lectures in noisy andreverberant environments. Index Terms : speaker localization, head orientation, microphonearrays, global coherence field. 1. Introduction Since 1990, several Speaker LOCalization (SLOC) techniqueshave been proposed as reported in [1, 2]. Most of the traditionalSLOC techniques are based on the estimation of time differencesof wavefront arrival at each sensor and on a consequent applica-tion of geometrical information to infer the acoustic source posi-tions. One of the most common techniques for Time Delay Es-timation (TDE) is based on Generalized Cross-Correlation PhaseTransform (GCC-PHAT) [3, 4]. Other effective SLOC techniquesare based on a preliminary computation of an acoustic map, asfor instance the Global Coherence Field (GCF) [5] representation,fromwhichthemostlikelysourcepositionisderivedthroughmax-imization in space.This paper aims at describing a new SLOC method that wasconceived starting from the effectiveness of the Oriented GlobalCoherence Field(OGCF), introduced in [6], which allows to char-acterize the orientation of an active speaker’s head with a satis-factory accuracy (in terms of angle error) even under reverberantconditions. By exploiting OGCF information, one can also derivemore robust speaker position estimates, since they are mostly re-lated to the propagation of a direct wavefront from a given point.On the other hand, previous SLOC techniques did not deal withthe way the sound is being radiated from a hypothesized positionin space.Theproposed method requires touseadistributed microphonenetwork similar to those available in the laboratories involved inthe EC CHIL

Read the paper · More papers on PaperTik