Speech Recognition Paradigms: A Comparative Evaluation of SpeechBrain, Whisper and Wav2Vec2 Models
Dhanvanth Reddy Yerramreddy, Jayasurya Marasani, Ponnuru Sathwik Venkata Gowtham, Guduru Harshit, Anjali Anjali · 2024
Speech recognition plays a pivotal role in the realm of natural language processing that deals in converting the language into the written text, providing human-computer interaction and enables us to use it widely for applications starting with voice assistants and delving upto the transcription services. Due to the complexity present in performing the task of speech recognition has led to the development of various models to enhance accuracy and efficiency. Our project mainly delves into three prominent speech recognition models that are Whisper, Wav2Vec2 and Speechbrain each of them representing distinct approaches of transcribing spoken language. The significance of the models used lies in their potential to perform better for realtime applications by improving the accuracy and reliability of speech recognition. To evaluate the best model that performs effectively, an array of metrics are used including levenshtein distance and it’s similarity percentage, jaccard similarity along with semantic similarity that provides an additional layer of evaluation describing the model’s understanding of contextual meaning in spoken language. Out of the three models, Speech Brain model outperformed all the other models through the calculation of Word Error Rate (WER), Character Error rate (CER) and BLEU score. The results have shown the model’s efficiency in converting spoken language(audio files) into precise and contextually relevant text. The results shows that these models contribute to the field of speech recognition highlighting the strengths of each approach and among them considering speechbrain as an ideal solution for accurate and meaningful transcription.