Optimizing Spectral Feature Based Text-Independent Speaker Recognition

Tomi Kinnunen · UEF eRepo (University of Eastern Finland) · 2005

Abstract A UTOMATIC speaker recognition has been an active research area for morethan 30 years, and the technology has gradually matured to a state ready forreal applications. In the early years, text-depended recognition was morestudied but gradually the focus has moved towards text-independent recognitionbecause their application fleld is much wider, including forensics, teleconferencing,and user interfaces in addition to security applications.Text-independentspeakerrecognitionis considerablymoredi–cultproblemcom-pared to text-depended recognition because the recognition system must be preparedfor an arbitrary input text. Commonly used acoustic features contain both linguisticand speaker information mixed in highly complex way over the frequency spectrum.The solution is to use either better features or better matching strategy, or a com-bination of the two. In this thesis, the subcomponents of text-independent speakerrecognition are studied, and several improvements are proposed for achieving betteraccuracy and faster processing.For feature extraction, a frame-adaptive fllterbank that utilizes rough phoneticinformation is proposed. Pseudo-phoneme templates are found using unsupervisedclustering, and frame labeling is performed via vector quantization, so there is noneed for annotated training data. For speaker modeling, experimental compari-iii

Read the paper · More papers on PaperTik