Unsupervised Speech-text word-level alignment with Dynamic Programming

Tianshu Yu, Zihan Gong, Minghuan Tan, Guhong Chen, Min Yang · 2025

Word-level alignment in speech-text pretraining has demonstrated significant effectiveness, particularly with models like SPECTRA that enhance cross-modal interactions and understand multi-turn dialog contexts.However, these advancements are constrained by a reliance on word-level annotated data, limiting their broader applicability and failing to fully exploit the vast amount of unannotated data available.This paper introduces an Unsupervised Speech-text word-level alignment with Dynamic Programming (USDP), which reduces the dependency on scarce annotated resources.We propose an iterative training method for USDP, inspired by the EM algorithm.This approach uses Dynamic Programming and EM principles to iteratively refine temporal alignment predictions.Initially, corresponding speech segments are identified based on the model's temporal predictions.A predictor then forecasts text words, and Dynamic Programming is applied to determine the optimal alignment, further refining the model's predictions.Furthermore, we conduct extensive experiments on six benchmark datasets across four different downstream speech-text tasks, including Emotion Recognition in Conversation (ERC), Multimodal Sentiment Analysis (MSA), Spoken Language Understanding (SLU), and Dialogue State Tracking (DST).The experimental results demonstrate that our method significantly enhances the accuracy of models on these speech-text downstream tasks compared to existing approaches.

Read the paper · More papers on PaperTik