Prosodic Modeling for Speaker Recognition Based on Sub-Band Energy Temporal Trajectories

André Adami · 2006

Recent work has proposed the use of a discrete representation of the dynamics of the fundamental frequency and short-term energy temporal trajectories to characterize speaker and/or language information. Since the short-term energy trajectory is affected by several factors, like speaker, phone, and channel information, we propose the use of the temporal trajectories from frequency bands instead of the short-term energy in the speaker modeling. This approach allows us to use only the relevant information (i.e., speaker and phone) and discard the irrelevant information (i.e., channel). The proposed approach is evaluated on the 2001 and 2003 NIST Extended-data speaker detection tasks. We show that the proposed approach can achieve 12% relative improvement in performance over the approach using the short-term energy trajectory.

Read the paper · More papers on PaperTik