Multitask Learning Model with Text and Speech Representation for Fine-Grained Speech Scoring
Seongjin Park, Rutuja Ubale · 2023
The goal of the present study is to evaluate whether computational models can accurately approximate human perceptual judgments of language learners’ proficiency. To do so, we develop an end-to-end multi-task model that predicts three sub-level proficiency scores (delivery, language use, and topic development) as well as a holistic score based on speech and text representations leveraging transformer-based architectures. We compared the performance of multi-task models with a baseline model and single-task models to examine the benefits of using transformer-based architectures and multi-task learning setting. Our results suggest that a multi-task setting and transformer-based models, particularly a bi-modal model, outperform baseline models in generalizing to unseen data. Furthermore, our findings indicate that speech may contain syntactic and semantic information that should be explored in future studies.