Deep Hashing for Speaker Identification and Retrieval Based on Auditory Sparse Representation
Dung Kim Tran, Masato Akagi, Masashi Unoki · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022
Deep hashing algorithms used for speaker identification and retrieval aim to produce discriminative hash codes for a set of speech signals. The basis of this task is highly related to speaker individuality. However, existing deep speaker hashing algorithms were constructed without considering speaker individuality. Previous studies have demonstrated the importance of speaker individuality in speech analysis and synthesis applications. Furthermore, recent studies have demonstrated the advantages of sparse representations of speech signals over the traditional spectrograms. This paper proposes a method for hashing the significant acoustical features related to speaker individuality by using auditory sparse representations. In speaker identification and speaker retrieval experiments with the VoxCeleb2 dataset, our 64-bit hash codes achieved 99.91% in top-1 accuracy and 97.55% in MAP@100, which are highly competitive with other state-of-the-art methods.