Modeling dynamic prosodic variation for speaker verification

Kemal Sönmez, Elizabeth E. Shriberg, Larry P. Heck, Mitchel Weintraub · 1998

Statistics of frame-level pitch have recently been used in speaker recognition systems with good results [1, 2, 3]. Although they convey useful long-term information about a speaker's distribution of f 0 values, such statistics fail to capture information about local dynamics in intonation that characterize an individual's speaking style. In this work, we take a first step toward capturing such suprasegmental patterns for automatic speaker verification. Specifically, we model the speaker's f 0 movements by fitting a piecewise linear model to the f 0 track to obtain a stylized f 0 contour. Parameters of the model are then used as statistical features for speaker verification. We report results on 1998 NIST speaker verification evaluation. Prosody modeling improves the verification performance of a cepstrum-based Gaussian mixture model system (as measured by a task-specific Bayes risk) by 10%. 1. INTRODUCTION Statistics of frame-level pitch have recently been shown to improve the perf...

Read the paper · More papers on PaperTik