Study for Automatic Speech Recognition for Wav2Vec2.0

Xiwei Huang · Applied and Computational Engineering · 2025

Automatic Speech Recognition (ASR) is a popular technology that converts speech audio into corresponding text. This application serves critical roles in areas such as virtual assistants, transcription services, accessibility tools, etc. This paper mainly introduces the application of the Wav2Vec2.0 model, which is an advanced self-supervised ASR model. The dataset used in this research is the Mozilla Common Voice dataset, which contains audio data in multiple languages and from people across different ages, genders, and occupations. In addition, the data preprocessing process and the architecture of the model will also be discussed in this research. The implementation demonstrates the strong ability of the Wav2Vec2.0 model in transcribing speech data from the Mozilla Common Voice database, and the experimental results highlight the model's robustness in handling variations in accent, speaking speed, and recording quality, achieving competitive word error rates (WER) across diverse linguistic scenarios. Results also indicate potential improvements in accuracy through more careful and targeted data processing and improving the tokenizer. All these findings underscore the model's future capacity in real-world speech recognition systems, emphasizing its adaptability and efficiency.

Read the paper · More papers on PaperTik