Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective

Hankun Wang, Haoran Wang, Yiwei Guo, Zhihan Li, Chenpeng Du, Kai Yu · 2026

Although text-based large language models exhibit human-level writing ability, end-to-end speech language models (SLMs) still struggle to generate semantically coherent outputs without explicit text transcription. There are several potential reasons for this performance degradation: (A) speech tokens mainly provide phonetic information rather than semantic information, (B) the length of speech sequences is much longer than that of text sequences, and (C) paralinguistic information, such as prosody and accent, introduces additional variability. In this paper, we explore the influence of three key factors separately by transitioning the modality from text to speech in an evolving manner. Our findings reveal that the impact of the three factors varies. Factor A has a relatively minor impact, factor B influences syntactical and semantic modeling more significantly, and factor C exerts the most substantial impact, particularly in basic lexical modeling. Based on these findings, we provide insights into the unique challenges of training SLMs and highlight pathways to develop more effective end-to-end SLMs.

Read the paper · More papers on PaperTik