Use of Generative AI for Audio Speech Recognition: Methods and Selection Criteria
Satish L. Parmanand, Pankaj H. Chandankhede, Atul R. Deshmukh · 2025
The generative AI has transformed human speech recognition systems audio-wise it was impossible prior. In this paper, various generative AI methods for speech recognition like End-to-end deep learning models, such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) and Transformer Models are described. A comprehensive evaluation framework is introduced in it to aid in the selection of appropriate generative AI models with respect to their performance and computational efficiency as well as language/dialect neutrality. Results indicate that, though generative AI substantially improves the performance of speech recognition in relation to accuracy and robustness, some guidelines need to be defined for an effective selection procedure aimed at maximizing system efficiency. Furthermore, the study highlights a call for further work to improve integration of these powerful AI models in practice-oriented solutions and establish their effectiveness and scalability for practical deployment. Improving these guidelines will allow for more accurate and trustworthy speech recognition systems going forward.