Oral Fluency Classification for Speech Assessment
Ashish Kumar Panda, Rajul Acharya, Sunil Kumar Kopparapu · 2023
Automatic speech quality assessment finds importance in evaluating the quality of spoken speech, especially by L2 speakers. Goodness of pronunciation, stressing on the right syllable in a multi-syllable word, and oral fluency are a few main components which are assessed for a speaker. While gauging pronunciation and identifying the syllable stress is relatively standard, oral fluency assessment has large variation in the rubrics used in addition to qualitative dimension to measure the quality of fluency. In this paper, we explore, using Statistical Machine Leaning (SML) and Deep Learning (DL) models, to classify oral fluency using two publicly available datasets, namely, Avalinguo Audio Dataset (AAD) and SpeechOcean762 (SO762). We introduce pre-trained DeepSpeech model embeddings in conjunction with known speech features like Mel- Frequency Cepstral Features (MFCC) and Fluency Features (FF) to correctly predict the fluency class. The best classification accuracy obtained for AAD was 95.04%, while the same for SO762 was 77.12%.