Speaker Detection Without Models

Dan Gillick, Seth Stafford, Barbara Peskin · 2006

In order to capture sequential information and to take advantage of extended training data conditions, we developed an algorithm for speaker detection that scores a test segment by comparing it directly to similar instances of that speech in the training data. This non-parametric technique, though at an early stage in its development, achieves error rates close to 1% on the NIST 2001 extended data task and performs extremely well in combination with a standard Gaussian mixture model system. We also present a new scoring method that significantly improves performance by capturing only positive evidence.

Read the paper · More papers on PaperTik