SAMU-XLSR: Semantically-Aligned Multimodal Utterance-Level Cross-Lingual Speech Representation
Sameer Khurana, Antoine Laurent, James Glass · IEEE Journal of Selected Topics in Signal Processing · 2022
We propose the ($\tt SAMU\text{-}XLSR$):Semantically-AlignedMultimodalUtterance-levelCross-LingualSpeechRepresentation learning framework. Unlike previous works on speech representation learning, which learns multilingual contextual speech embedding at the resolution of an acoustic frame (10–20 ms), this work focuses on learning multimodal (speech-text) multilingual speech embedding at the resolution of a sentence (5–10 s) such that the embedding vector space is semantically aligned across different languages. We combine state-of-the-art multilingual acoustic frame-level speech representation learning model$\tt XLSR$with the Language Agnostic BERT Sentence Embedding ($\tt LaBSE$) model to create an utterance-level multimodal multilingual speech encoder$\tt SAMU\text{-}XLSR$. Although we train$\tt SAMU\text{-}XLSR$with only multilingual transcribed speech data, cross-lingual speech-text and speech-speech associations emerge in its learned representation space. To substantiate our claims, we use$\tt SAMU\text{-}XLSR$speech encoder in combination with a pre-trained$\tt LaBSE$text sentence encoder for cross-lingual speech-to-text translation retrieval, and$\tt SAMU\text{-}XLSR$alone for cross-lingual speech-to-speech translation retrieval. We highlight these applications by performing several cross-lingual text and speech translation retrieval tasks across several datasets.