Detection of AI Generated Speech using Speech Recognition with MFCC and GMM
Abhinav Sharma, Anshu Sharma, Utkarsh Pant · 2024
Generative AI-based Speech is nearly indistinguishable for a human ear. It can be misused for getting fraudulent access to a secured system, passing the wrong information or orders, committing financial frauds etc. The presented work aims to develop an AI generated speech detection system with the use of speech recognition technique for classification of AI generated speech (AI) and an authenticated human speech (H). In this work, AI Speaker Recognition has been performed in English language by the extraction of Mel Frequency Cepstral Coefficients (MFCCs). Gaussian Mixture Models (GMM) technique has been used for Classification. the human (noiseless) speech utterances and their AI generated files were used for training the GMM’s for two classes. The validation has been implemented using a large dataset of (i) Human (ii) the concatenated versions of Human of both classes (human and AI generated) with multiples of different noises such as fan sound, sound of crowd-talking, general traffic noise, light/medium and heavy rain sounds, sound on the construction site etc. (iii) the random speech samples recorded and created at multiple public places (both human and AI generated), to validate the robustness of the detection system. The recognition accuracy has been tested by varying the number of MFCCs extracted from small frames of a speech sample. More the number of (i) MFCC extracted and/or (ii) the number of training iterations and/or (iii) the number of gausses of GMM causes more recognition accuracy to a certain level but more computation time is elapsed too. An optimum combination of number of features, elapsed computation time, and accuracy is achieved. Results show that accuracy is highest (96 %) at 13 numbers of MFCCs.