Enhancing Communication: Utilizing Transfer Learning for Improved Speech-to-Text Transcription

D. Sasikala, Shaik Huzaifa Fazil · 2024

Automatic Speech Recognition (ASR) transforms spoken language into text facilitating interaction with technology and easing access to information for individuals facing challenges in communication through conventional text-based methods due to conditions like mobility impairments or speech impairments. The objective of this work is to develop a custom speech-to-text system that aids people with disabilities. In this study, transfer learning approach is explored with wav 2 vec a pretrained model. Fine-tuned wav2vec on the TIMIT dataset achieved a Word Error Rate (WER) of 30%, demonstrating the merits of transfer learning. In summary, this work reveals that transfer learning with pre-trained models like wav2vec 2.0 can enable accurate ASR for assistive technology applications with limited training data. This work aims to leverage speech-to-text transcription more accessible for people with disabilities.

Read the paper · More papers on PaperTik