Optimizing Speech Recognition for Medical Transcription: Fine-Tuning Whisper and Developing a Web Application

Rakesh Roushan, Harshit Mishra, Lucky Yadav, Sreeja Koppula, Nitya Tiwari, K. S. Nataraj · 2024

Utilizing automatic speech recognition in medical transcription can greatly improve healthcare professionals’ productivity. It eliminates the need for manual tasks like physician note-taking, data retrieval, and medical information searches, which can be time-consuming and divert their attention from patient care. Fine-tuning automatic speech recognition (ASR) systems for medical transcription using domain specific data is essential to enhance the performance. In this work, we fine-tuned the Whisper ASR system, which is known for its state-of-the-art speech recognition capabilities, using medical speech data. The fine-tuned model achieved a word error rate (WER) of 7.5% for medical data, demonstrating its potential for accurate transcription in clinical settings. Futher, we have developed a web application tailored for medical transcription. This application uses the robust Whisper ASR engine, renowned for its resilience to background noise. The web application offers user-friendly features such as audio recording, secure access via user-specific logins, and the seamless preservation of medical speech reports. We tested the web application with six Indian doctors by recording 362 utterances of prescriptions. The speech transcription achieved a WER of 19.4%, indicating the need for further fine-tuning for the Indian context with a larger dataset and an advanced Whisper model variant.

Read the paper · More papers on PaperTik