A ROBUST UNSUPERVISED WORD BY WORD TRANSLATION FOR MORPHOLOGICAL RICH LANGUAGES USING DIFFERENTRETRIEVAL TECHNIQUES

Shweta Chauhan, Umesh Pant, Mustafa, Philemon Daniel · Journal of Critical Reviews · 2020

Abstract Word translation or incorporation of bilingual dictionaries is an important capability that impacts many multilingual language processing tasks. In recent years, cross-lingual word embedding has been receiving considerable attention. Recently, it has been shown that these word embeddings can be learned by aligning two monolingual disjoint vector spaces via linear transformations, using as supervision no more than a small bilingual dictionary. In this paper, the best cross-lingual word embedding is generated for English as source language, and Hindi, Punjabi, Telegu and Tamil are target language and vice versa. Here, there is no aligned document or sentence aligned corpus, nor any bilingual dictionary has been considering. We are following the assumption of intralingual similarity distribution that, for the most common word, the distribution graph is similar between language pair and embeddings are isometric. Different types of retrieval methods nearest neighbor, inverted nearest neighbor retrieval, inverted Softmax, and cross-lingual word scaling are performed and compared for the bi-lingual embedding of language pairs, which is trained for fully unsupervised learning techniques. Bi-lingual word embedding is tested on generated English-Hindi, English-Punjabi, English-Telegu English-Tamil dictionary..

Read the paper · More papers on PaperTik