Potential scope of a fully-integrated architecture for speech translation

Alicia Pérez, María Inés Torres, Francisco Casacuberta · 2010

The classical approach to tackle speech translation assembles a text-to-text translation system placed after a speech recogniser, yielding the so-called decoupled architecture. In this regard, there are two issues to bear in mind: first, what is translated in the decoupled architecture is the most likely transcription of the spoken utterance; second, translation systems are sensitive to errors in the source string, and speech recognition systems are still far from being flawless. In this paper we promote the use of an architecture to carry out speech translation that allows to build up the most likely translation relying upon both acoustic and translation models in a cooperative manner, that is the so-called integrated architecture. The integrated architecture is implemented in the finite-state framework by virtue of the composition of finite-state acoustic models of the source language within a stochastic finite-state transducer that would encompass source and target languages. The potential performance of the integrated architecture is assessed quantitatively in relation to the decoupled one. We conclude that while the single-best approach for both decoupled and integrated architectures show similar performance, an oracle evaluation reveals that the potential scope of the integrated architecture would offer statistically significant differences. c ○ 2010 European Association for Machine Translation. 1 Statistical speech translation The goal of statistical speech translation is to seek the most likely string in the target language, ̂t, given the acoustic representation of a speech signal in the source language, x. ̂t = argmax P (t|x) (1) t The source string, s, that is the transcription of the speech utterance x, can be introduced as a hidden variable the Bayes ’ decision rule applied (Ney,

Read the paper · More papers on PaperTik