Exploiting parallel corpus for automatic extraction of multilingual names: Transliteration perspective
Bibekananda Kundu, Sanjay Kumar Choudhury · 2012
This paper describes a novel approach for extraction of multilingual transliteration pairs from aligned parallel corpus. The proposed approach utilizes an encoding technique based on “Place and Manner of Articulation”. Jaccard Coefficient has been used to measure the distance between encoded source and target transliteration pairs. The proposed methodology has been employed for extraction of English-Bangla transliteration pairs and reported 94% accuracy which is quite encouraging when compared to Expectation Maximization based word alignment module that yields 59.42% accuracy on the same test data.