Automatic speaker recognition: Advances toward informative systems.

Daniel Garcia-Romero, Carol Espy-Wilson · The Journal of the Acoustical Society of America · 2010

Joint factor analysis (JFA) has become the state-of-the-art in automatic speaker recognition systems. In this paradigm, the information contained in a variable-length speech recording is summarized as a fixed-length supervector by means of a soft partition of the acoustic space through a Gaussian mixture model. Moreover, an explicit mechanism to account for the speaker and undesired (inter-session) variability in terms of a small set of factors results in very accurate answers to the question of whether two speech samples are uttered by the same speaker or not. However, no apparent answer to the question of what exactly is that makes two particular voices similar or different is obtained by this approach. To address this issue of interpretability, we propose the modification of the standard JFA approach in two ways. First, by explicitly incorporating phonetic information in the construction of the supervectors so that subsets of its entries get associated with specific phonetic contexts. Second, by modifying the estimation of the speaker and inter-session factors so that phonetically contextualized factors are obtained. A study of the recognition accuracy as well as the interpretability of the results of the proposed approach will be performed on the NIST 2008 speaker recognition evaluation.

Read the paper · More papers on PaperTik