Identifying Similar and Co-referring Documents Across Languages

Pattabhi R. K. Rao, Sobha Lalitha Devi · International Joint Conference on Natural Language Processing · 2008

This paper presents a methodology for finding similarity and co-reference of documents across languages. The similarity between the documents is identified according to the content of the whole document and co-referencing of documents is found by taking the named entities present in the document. Here we use Vector Space Model (VSM) for identifying both similarity and co-reference. This can be applied in cross-lingual search engines where users get documents of very similar content from different language documents.

Read the paper · More papers on PaperTik