Streaming Arabic: Bridging The Gap Between Offline and Real-Time ASR Performance
Mohamed Maged Abdelazeem Elsetohy · 2026
Streaming automatic speech recognition (ASR) is critical for real-time applications, while encoder--decoder models have recently achieved state-of-the-art performance in offline ASR. Traditionally, streaming ASR has relied on encoder-only architectures, such as CTC- and transducer-based models, but recent research has begun to explore how the modeling strengths of encoder--decoder architectures can be brought into low-latency settings. This problem remains largely unexplored for compact language-specific ASR models, especially for Arabic. In this thesis, inspired by recent progress in simultaneous speech translation, we used the cross-attention of a state-of-the-art attention-based Arabic ASR model to enable streaming recognition. We show that directly applying a streaming policy results in a substantial degradation relative to offline decoding, with a 4\% to 10\% increase in word error rate (WER). Through detailed analysis, we find that this degradation is primarily caused by insertion errors in short input segments. This issue appears to be more pronounced for language-specific models trained on comparatively limited amounts of data, while it is less severe for large multilingual models such as Whisper. To mitigate this problem, we propose a lightweight length adaptation method that reduces insertion errors and narrows the offline--streaming gap to approximately 1\% WER, achieving state-of-the-art performance in streaming Arabic ASR. We also extend the proposed approach to five Arabic dialects and show that length adaptation remains beneficial even when applied to a different ArTST variant.