Discriminative Power of Transient Frames in Speaker Recognition
Jérôme Louradour, Khalid Daoudi, Régine Andre-Obrecht, Paul Sabatier · 2006
In speaker recognition, several recent studies attempt to integrate prior knowledge in order to make better distinction between speakers. This paper studies the relative speaker discriminative power of speech transient and steady zones. An automatic segmentation is used to localize these two types of zones. Experiments are carried out with the KIST 2003 speaker evaluation database. They show that transient frames, in the neighborhood of segment boundaries, are more speaker discriminative than the middle frames of long segments which correspond to the steady parts of phones. In addition, this study leads to a higher efficiency system (than the baseline one) by processing roughly three times less data during the scoring stage.