Noise robust speaker recognition with convolutive sparse coding
Antti Hurmalainen, Rahim Saeidi, Tuomas I. Virtanen · 2015
Recognition and classification of speech content in everyday environments is challenging due to the large diversity of real-world noise sources, which may also include competing speech. At signal-to-noise ratios below 0 dB, a majority of features may become corrupted, severely degrading the performance of clas-sifiers built upon clean observations of a target class. As the energy and complexity of competing sources increase, their ex-plicit modelling becomes integral for successful detection and classification of target speech. We have previously demon-strated how non-negative compositional modelling in a spec-trogram space is suitable for robust recognition of speech and speakers even at low SNRs. In this work, the sparse coding approach is extended to cover the whole separation and clas-sification chain to recognise the speaker of short utterances in difficult noise environments. A convolutive matrix factorisation and coding system is evaluated on 2nd CHiME Track 1 data. Over 98 % average speaker recognition accuracy is achieved for shorter than three second utterances at +9...-6 dB SNR, illus-trating the system’s performance in challenging conditions. Index Terms: speaker recognition, noise robustness, composi-tional models, sparse coding, non-negative matrix factorization