Mining Transliterations from Wikipedia Using Pair HMMs
Peter Nabende · 2010
This paper describes the use of a pair Hidden Markov Model (pair HMM) system in mining transliteration pairs from noisy Wikipedia data. A pair HMM variant that uses nine transition parameters, and emission parameters associated with single character mappings between source and target language alphabets is identified and used in estimating transliteration similarity. The system resulted in a precision of 78 % and recall of 83 % when evaluated on a random selection of English-Russian Wikipedia topics. 1