Using deep belief networks for vector-based speaker recognition

William M. Campbell · 2014

Deep belief networks (DBNs) have become a successful ap-proach for acoustic modeling in speech recognition. DBNs ex-hibit strong approximation properties, improved performance, and are parameter efficient. In this work, we propose meth-ods for applying DBNs to speaker recognition. In contrast to prior work, our approach to DBNs for speaker recognition starts at the acoustic modeling layer. We use sparse-output DBNs trained with both unsupervised and supervised methods to gen-erate statistics for use in standard vector-based speaker recogni-tion methods. We show that a DBN can replace a GMM UBM in this processing. Methods, qualitative analysis, and results are given on a NIST SRE 2012 task. Overall, our results show that DBNs show competitive performance to modern approaches in an initial implementation of our framework. Index Terms: speaker recognition, deep belief networks 1.

Read the paper · More papers on PaperTik