Feature Fusion for Performance Enhancement of Text Independent Speaker Identification

Zahra Shah, Giljin Jang, Adil Farooq · ICCK Transactions on Intelligent Systematics · 2024

Speaker identification systems have gained significant attention due to their potential applications in security and personalized systems. This study evaluates the performance of various time- and frequency-domain physical features for text-independent speaker identification. Four key features—pitch (P), intensity (I), spectral flux (SF), and spectral slope (SS)—were examined along with their statistical variations (minimum, maximum, and average values). These features were fused with log power spectral features and trained using a Convolutional Neural Network (CNN). The goal was to identify the most effective feature combinations for improving speaker identification accuracy. The experimental results revealed that the proposed feature fusion method outperformed the baseline system by approximately 8%, achieving an accuracy of 87.18%.

Read the paper · More papers on PaperTik