Unveiling the Fundamental Obstacle in Speech-to-Text Modeling: Understanding and Mitigating the Granularity Challenge

Chen Xu, Xiaoqian Liu, Yuhao Zhang, Anxiang Ma, Tong Xiao, Jingbo Zhu, Dapeng Man, Wu Yang · IEEE Transactions on Audio Speech and Language Processing · 2025

Speech-to-text (S2T) generation tasks often struggle to achieve satisfactory convergence without relying on auxiliary data or models. We identify the core issue as the modeling granularity, where the fine-grained and lengthy characteristics of audio features pose obstacles in effectively allocating attention weights, particularly during encoder self-attention learning. In this paper, we investigate two well-established methods, Conformer and information aggregation, to mitigate the learning burden of the encoder from the aspects of intra-layer and inter-layer encoding. Conformer directly enhances modeling capability through architecture improvement, while aggregation generates coarser-grained representations, thus shaping text-like structures to fundamentally simplify attention learning. Extensive results demonstrate superior convergence and notable improvement on two representative S2T generation tasks, speech recognition and translation. In particular, we achieve a notable average BLEU score of 26.9 on MuST-C speech translation datasets without auxiliary resources and approaches. Furthermore, our finding suggests that the grown model capacity and sufficient training can effectively facilitate the granularity challenge.

Read the paper · More papers on PaperTik