Early Fusion of Phone Embeddings for Recognition of Low-Resourced Accented Speech
Sunakshi Mehra, Seba Susan · 2022
We describe a supervised technique in this paper for accented speech recognition using constrained resources. One problem with existing research is that less work has been accomplished on using phonology for understanding spoken text. We suggest using early fusion of phone embeddings to recognize accented speech from a small sample dataset. Using PocketSphinx, the phonemes are recovered from the .WAV recordings. FastText's character n-gram based subword embeddings transforms the phonemes into vectors. To keep the vectors uniform, we concatenated and padded the vectors. The accuracy of the early fusion of phone embeddings for accented speech recognition is 49.01% for 10 sentence categories of the L2- ARCTIC accented speech corpus, which is higher than the accuracy of existing techniques. We seek to prove through our work that audio phonemes can play a significant role in accented speech recognition even when the number of training samples are few.