Development of text and speech corpus for an Indonesian speech-to-speech translation system

Mohammad Teduh Uliniansyah, Hammam Riza, Agung Santosa, Gunarso, Made Agus Oka Gunawan, Elvira Nurfadhilah · 2017

This paper describes our natural language resources especially text and speech corpora for developing an Indonesian speech-to-speech translation (S2ST) system. The corpora are used to create models for Automatic Speech Recognition (ASR), Statistical Machine Translation (SMT), and Text-to-Speech (TTS) systems. The corpora collected since 1987 from various sources and projects such as Multilingual Machine Translation System (MMTS), PAN Localization, ASEAN MT, U-STAR, etc. Text corpora are created by either collecting from online resources or translating manually from textual sources. Speech corpora are made from several recording projects. Availability of these corpora enables us to develop Indonesian speech-to- speech translation system.

Read the paper · More papers on PaperTik