Speech Recognition for Composite Languages with Low-Density Components
Joshua Waxman, Rachel Schachter, Gila Linzer, Chaya Trapedo, Deen Fink, Ariel Melnitsky · 2024
We present an approach to automatic speech recognition for hybrid languages with frequent code switches, low density component languages, and dialectal differences in pronunciation. Whisper, a multitask and multilingual speech model, performs automatic speech recognition quite well for many high-resource languages but less so for low-resource languages. We are targeting a language, Yeshivish, with features that make recognition difficult: it is hybrid, with frequent code switches, low density component languages, and dialectal differences in pronunciation. Our approach to improving the accuracy of Whisper transcription for low-resource hybrid languages involves an initial pass using the primary language model; a correction step using ChatGPT or word-alignment with a known text; and fine-tuning Whisper with the partially corrected data.