Integrating ASR and RoBERTa for Automated Oral English Proficiency Scoring

Soma Sircar Dasgupta, Yaisna Rajkumari, Freeda Rajakumari, K.M. Madhumitha, T. Jayasudha, B Sreela · 2025

The study explores the integration of Automatic Speech Recognition (ASR) technology and RoBERTa, a transformer-based language model, for assessing oral English proficiency in language learners. Traditional methods of evaluating pronunciation often rely on human assessors, which can be subjective, time-consuming, and lack scalability. Additionally, these methods struggle to provide instant, individualized feedback to learners. The proposed work aims to address these limitations by developing an automated oral proficiency scoring system that combines ASR and RoBERTa for real-time, objective evaluation of pronunciation accuracy, fluency, intonation, and stress. The objective of this study is to assess the effectiveness of ASR-RoBERTa in providing reliable, real-time feedback on English pronunciation, focusing on aspects such as phoneme-level accuracy, speech fluency, and prosodic features. The novelty of this approach lies in the fusion of ASR's acoustic feature extraction with RoBERTa's contextual language modeling for more comprehensive evaluation. Findings indicate that the ASR-RoBERTa system closely aligns with human rater scores, with strong correlations for pronunciation accuracy (0.92), fluency (0.89), intonation and stress (0.87), and overall proficiency (0.91). The system's automated nature offers scalable, immediate feedback to learners, demonstrating its potential in enhancing language assessment and learning outcomes.

Read the paper · More papers on PaperTik