Using Deep Learning Neural Networks to Recognize and Authenticate the Identity of the Speaker

Seddiq Q. Abd Al-Rahman, Sameeh Abdulghafour Jassim · 2024

Nowadays, the advancement of technology is given that permeates all parts of daily life and is dependent on electronic services like banking and financial transfers, project management, health care, and other essential areas. These applications' fundamental components (person recognition and/or authentication procedures) can be viewed as one of their challenging constraints. Therefore, using biometric features in these sectors can result in promising outcomes. A person's voice is a distinctive bio-feature that may be used for identification, verification, and prevents others from mistaking them for someone else without their knowledge or consent. The Speaker Recognition System that was proposed utilized the biometric method to confirm the speaker's identity and then authenticate that identity by using the password to prevent replay attacks. The objective of the system is to distinguish the speaker by utilizing the corpus of the English language dataset of 20 adult speakers (12 male and 8 female) and for each person there 500 utterances samples were utilized for the training and testing purpose equally. Additionally, the input speech signal is processed to get a specific number of variables, named features. Features extraction is utilized to select the most appropriate features to reduce the computational power load and thus reduce the time needed for speaker recognition. MFCC and LPCCs are integrated as speech feature extraction and DFFNNs as speech classifiers. The classification classifies the speaker by utilizing the features defined by the feature extraction block and audio clips in the database. Furthermore, the MATLAB (R2023a) tool was utilized for the entire simulation. The result got from this proposal proved that the integration of MFCC and LPCC techniques has higher accuracy than utilizing every technique individually as well as utilizing the DFFNNs to classify the speakers' got high recognition results among the speakers. As well as the authentication process will be more accurate because we used two authentication factors (What you know(password) and What you are (speaker recognition).

Read the paper · More papers on PaperTik