Temporal-Spectral Analysis for Speaker Identification and Authentication in Emotional Speech
Mohamed Alae-Eddine Eladlani, Larbi Boubchir, Khadidja Benallou · 2024
Automatic speaker identification and verification in diverse acoustic environments remains a major challenge due to the high variability of speech signals, particularly due to emotional variations and ambient noise. This paper presents a novel approach to improve the robustness and accuracy of speaker recognition systems in these complex conditions. Our method combines advanced preprocessing of audio signals with a hybrid neural network architecture, designed to efficiently capture the spatiotemporal characteristics of speech spectrograms. The proposed preprocessing process transforms audio signals of varying durations into uniform spectral representations, facilitating their analysis by deep learning models. The developed neural architecture integrates Temporal Convolutional Layers (TCN) and Long Short-Term Memory (LSTM) recurrent networks, enriched by an additive attention mechanism exploiting auxiliary information on gender and emotions. Experiments conducted on the RAVDESS and CREMA-D databases demonstrate the remarkable effectiveness of our approach. For speaker identification, the model achieves an accuracy higher than 96% on both bases. For speaker authentication, it displays a global Equal Error Rates (EER) of 1.75% for RAVDESS and 1.60% for CREMA-D, significantly outperforming some existing methods.