Modeling NERFs for speaker recognition.

Sachin S. Kajarekar, Luciana Ferrer, Kemal Sönmez, Jing Wen Zheng, Elizabeth E. Shriberg, Andreas Stolcke · 2004

We introduce a new type of feature to capture long-range patterns associated with individual speakers or with speaking styles. NERFs, or Nonuniform Extraction Region Features, are defined based on regions of speech that are delimited by various automatically extractable events of interest. There is a wide unexplored space of potentially useful NERFs, but to use them successfully, at least two important challenges must be addressed: (1) methods for coping with inherently missing features, and (2) methods for feature selection from large sets of potentially correlated NERFs. We address the issue of missing features in this paper. We propose three methods for modeling NERFs that cope with missing features. We show that on the 2003 NIST extended-data speaker recognition evaluation task, a NERF system yields an EER of 11.6% alone, and improves the MFCC baseline performance by roughly 15 % relative. within that region (the Fs). A region refers to a contiguous stretch of speech bounded by automatically extractable events of interest. A region could be bounded, for example, by pauses, by unstressed syllables, by pitch rises or falls, and so on. For defining potential NERs, we consider both what might constitute meaningful or characteristic units at some level of production (for example, a prosodic phrase) and what types of boundary events we can use to automatically delimit those units. The Fs are the features within the NERs. They can be defined to measure, for example, the maximum or mean pitch values, duration patterns, energy contours, and so forth. These features are similar to those used in studies of other unit types, such as utterances and words. 1.

Read the paper · More papers on PaperTik