Design of Scalability Enabled Low Cost Automatic Speaker Recognition System Using Light Weight Multistage Classifier
Husham Ibadi, S. S. Mande, Imdad Rizvi · IETE Journal of Research · 2023
Protracted learning processes and dedicated computational power are the quality determinants in automatic speaker recognition (ASR). Conventionally, speakers are separately modelled using their acoustic features and surrounding environment features. Analytical speaker modelling via the Gaussian mixture model (GMM) and the Universal Background Model (UBM) is predominant over invariant speech condition models. A high number of speakers, fluctuated surroundings, and limited capacity processors are persisting performance barriers in both analytical and automatic ASR models. This paper argues on the ASR scalability approach which ensures fine model performance under a large number of speakers with varying biometrics. It also aimed to bridge performance-hindering barriers via using Light Weight Multistage classification. Consequently, 100% recognition accuracy could be achieved over three different databases of 12, 19, and 118 speakers.