Hindi CLIR in thirty days

Leah S. Larkey, Margaret E. Connell, Nasreen AbdulJaleel · ACM Transactions on Asian Language Information Processing · 2003

As participants in the TIDES Surprise language exercise, researchers at the University of Massachusetts helped collect Hindi--English resources and developed a cross-language information retrieval system. Components included normalization, stop-word removal, transliteration, structured query translation, and language modeling using a probabilistic dictionary derived from a parallel corpus. Existing technology was successfully applied to Hindi. The biggest stumbling blocks were collection of parallel English and Hindi text and dealing with numerous proprietary encodings.

Read the paper · More papers on PaperTik