EBMT, SMT, hybrid and more: ATR spoken language translation system.
Eiichiro Sumita, Yasuhiro Akiba, Takao Doi, Andrew Finch, Kenji Imamura, Hideo Okuma, Michael D. Paul, Mitsuo Shimohata, Taro Watanabe · 2004
This paper introduces ATR’s project named Corpus-Centered Computation (C3), which aims at developing a translation technology suitable for spoken language translation. C3 places corpora at the center of its technology. Trans-lation knowledge is extracted from corpora, translation qual-ity is gauged by referring to corpora, the best translation among multiple-engine outputs is selected based on corpora, and the corpora themselves are paraphrased or filtered by automated processes to improve the data quality on which translation engines are based. In particular, this paper reports the hybridization archi-tecture of different machine translation systems, our tech-nologies, their performance on the IWSLT04 task, and para-phrasing methods. 1.