Dilated Convolution and MelSpectrum for Speaker Identification using Simple Deep Network

Hema Kumar Pentapati, K N Sridevi · 2022 8th International Conference on Advanced Computing and Communication Systems (ICACCS) · 2022

A challenge in the Speaker Recognition systems is variations in the speech utterances of a speaker in different instances. We propose a deep learning approach to develop the simple Convolution Neural Network(CNN) with limited layers and reduced data to train the network. The commonly used Log-MelSpectrum is employed to represent the speech signal. This network uses dilated convolution layer instead of traditional layer with deep stride and filters with increasing order. Dilated Convolution exhibits more receptive area which helps to hold the low level features in speech signal. The performance of the proposed network with clean Libri speech data and 200 samples of each speaker is compared against existing scheme. The accuracy reaches 76.1% using existing method, while it is increased to 80.6% when using the proposed method and without introducing greater time complexity.

Read the paper · More papers on PaperTik