On estimation of a speaker's confusion matrix from sparse data

Stephen Cox · 2008

Confusion matrices have been widely used to increase the ac-curacy of speech recognisers, but usually a mean confusion ma-trix, averaged over many speakers, is used. However, analysis shows that confusion matrices for individual speakers vary con-siderably, and so there is benefit in obtaining estimates of con-fusion matrices for individual speakers. Unfortunately, there is rarely enough data to make reliable estimates. We present a technique for estimating the elements of a speaker’s confu-sion matrix given only sparse data from the speaker. It utilizes non-negative matrix factorisation to find structure within confu-sion matrices, and this structure is exploited to make improved estimates. Results show that under certain conditions, this tech-nique can give estimates that are as good as those obtained with twice the number of utterances available from the speaker. 1.

Read the paper · More papers on PaperTik