Unsupervised Arabic Speech Embedding Model for Speaker Identification

Noora Al Roken, Abir Jaafar Hussain, Ismail Mohd Adnan Shahin, Ayad Mashaan Turky, Bilal Khan, Wasiq Khan · 2023

Speech modality has recently been explored at a large scale with the advancement of computational capabilities, forensics applications, and the availability of big data. Automation of speaker identification via advanced machine learning approaches is gaining momentum to improve various applications, particularly forensics and information security. Due to the lack of large Arabic datasets, existing speaker identification models are trained on small amounts of speech inputs which is inefficient when applied in real-life. In the present study, an unsupervised Arabic speech embedding model is created, leveraging large amounts of unlabeled speech data to help streamline the identification process for various security areas. The proposed framework is divided into two sections: unsupervised speech feature extraction and supervised speaker identification. An unlabeled Arabic speech dataset was constructed with approximately 1 million samples to develop the deep feature embedding model. Various machine learning classifiers and a deep neural network were used for speaker identification on the labeled Emirati speech database, which contains utterances of eight traditional Emirati phrases. The K-nearest neighbor classifier achieved the best result with the speech embeddings for speaker identification with an average accuracy of 93.58% using five-fold cross-validation.

Read the paper · More papers on PaperTik