Reinforced temporal structure information for embedded utterance-based speaker recognition
Anthony Larcher, Jean-François Bonastre, John S. D. Mason · 2008
Embedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of com-puting resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a rel-atively large quantity of speech data, but with negligible har-nessing of linguistic content. The proposed approach addresses these limitations and takes advantage from the linguistic nature of the speech material into the GMM/UBM framework by us-ing client-customised utterances. The GMM/UBM is then rein-forced with new temporal information. Experiments on the MyIdea database are performed when im-postors know the client-utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 45 % in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations. 1.