Data representation methods and use of mined corpora for Indian language transliteration

Anoop Kunchukuttan, Pushpak Bhattacharyya · 2015

Our NEWS 2015 shared task submission is a PBSMT based transliteration system with the following corpus preprocessing enhancements: (i) addition of wordboundary markers, and (ii) languageindependent, overlapping character segmentation.We show that the addition of word-boundary markers improves transliteration accuracy substantially, whereas our overlapping segmentation shows promise in our preliminary analysis.We also compare transliteration systems trained using manually created corpora with the ones mined from parallel translation corpus for English to Indian language pairs.We identify the major errors in English to Indian language transliterations by analyzing heat maps of confusion matrices.

Read the paper · More papers on PaperTik