Text-Independent Speaker Verification Using Lightweight 3D Convolutional Neural Networks

Jyun-Yan Chen, Jin-Tsong Jeng · 2024

This paper presents a text-independent speaker verification system leveraging lightweight 3D Convolutional Neural Networks (3D-CNN). Our system independently operates of text and focuses on classifying speakers based on their extracted features. We employ lightweight 3D-CNN to capture the nuances within speech samples from the same speaker. Initially, speaker speech data is used for enrollment, generating corresponding speaker features that form the basis of the speaker model, also referred to as the identity discriminator. Subsequently, speech data requiring verification is utilized as evaluation data, and the resulting speaker features are compared with the speaker model using cosine similarity. Experimental findings show that our system achieves 14.3% on Equal Error Rate (EER). Additionally, the performance of the lightweight 3D-CNN system remains consistent compared to the 3D-CNN system. At the same time, the proposed highlighting system can effective to reduce on computational cost.

Read the paper · More papers on PaperTik