Compensating forMismatch inHigh-Level Speaker Recognitiont

William M. Campbell · 2006

Speaker recognition using high-level features hasbeenasuccessful areaofexploration. Features obtained frommanydifferent levels phones, words, prosodic events, etc.areused tocharacterize thespeaker. A goodmodeling technique for these features isthesupport vector machine (SVM).SVMs modelthen-gramfrequencies fromspeaker utterances ina high-dimensional SVM feature space andhaveshownexcellent performance overawidevariety ofhigh-level features. Acomplimentary method ofrecent exploration inSVM speaker recognition istheuseofnuisance attribute projection (NAP). NAPremovesdirections fromSVM feature space that aresuperfluous tothetask ofspeaker recognition channel information, sessionvariability, etc.Inthis paper, weconsider theapplication ofNAPtohigh-level speaker recognition. Wedescribe thedifficulties inapplying this method andpropose solutions. Wealso conduct experiments showing that NAP canreduce variability inSVM feature space leading toimproved performance.

Read the paper · More papers on PaperTik