Aligning Speech-Text Representations via Contrastive Modality Translation

Jen‐Tzung Chien, Pin-Yen Liu · IEEE Transactions on Audio Speech and Language Processing · 2025

Recent advances in automatic speech recognition (ASR) have led to substantial improvements in system accuracy and robustness, particularly in converting speech signals into text sequences. Among prevailing ASR paradigms, both acoustic encoder architectures and end-to-end encoder-decoder frameworks have demonstrated strong performance across a variety of downstream tasks. However, challenges remain in achieving robust generalization, especially for the models trained solely on acoustic inputs. These limitations are exacerbated in scenarios involving complex linguistic content and variable acoustic conditions. To mitigate these issues, recent research has explored joint modeling of speech and text, aiming to embed both modalities into a shared representation space. While promising, such approaches often suffer from modality interference and suboptimal alignment, particularly under domain shifts. Notably, most existing evaluations have been conducted on neutral speech, with limited attention to robustness under emotional or affective speech conditions. This study proposes a contrastive modality translation framework to improve the alignment between speech and text representations within a unified embedding space. By leveraging cross-modal contrastive learning, the proposed method enhances ASR performance, especially in emotion-influenced speech scenarios. Empirical results demonstrate that our approach significantly improves recognition accuracy across both neutral and emotionally expressive speech, highlighting its effectiveness in addressing cross-modality and domain-shift challenges in ASR.

Read the paper · More papers on PaperTik