Leveraging Statistical Transliteration for Dictionary-Based English-Bengali CLIR of OCR'd Text
Utpal Garain, Arjun Das, David Doermann, Douglas W. Oard · International Conference on Computational Linguistics · 2012
This paper describes experiments with transliteration of out- of-vocabulary English terms into Bengali to improve the effectiveness of English-Bengali Cross-L anguage Information Retrieval. We use a statistical translation model as a basis for transliteration, and present evaluatio n results on the FIRE 2011 RISOT Bengali test collection. I ncorporating transliteration is shown to substantially and statistically significantly improve Mean Average Precision for both the text and OCR conditions. Learning a distortion model for OCR errors and then using that model to improve recall is also shown to yield a further substantial and statistically significant improvement for the OCR condition.