Multitask Learning Model with Text and Speech Representation for Fine-Grained Speech Scoring

Seongjin Park, Rutuja Ubale · 2023

The goal of the present study is to evaluate whether computational models can accurately approximate human perceptual judgments of language learners’ proficiency. To do so, we develop an end-to-end multi-task model that predicts three sub-level proficiency scores (delivery, language use, and topic development) as well as a holistic score based on speech and text representations leveraging transformer-based architectures. We compared the performance of multi-task models with a baseline model and single-task models to examine the benefits of using transformer-based architectures and multi-task learning setting. Our results suggest that a multi-task setting and transformer-based models, particularly a bi-modal model, outperform baseline models in generalizing to unseen data. Furthermore, our findings indicate that speech may contain syntactic and semantic information that should be explored in future studies.

Read the paper · More papers on PaperTik