Identifying Similar and Co-referring Documents Across Languages
Pattabhi R. K. Rao, Sobha Lalitha Devi · International Joint Conference on Natural Language Processing · 2008
This paper presents a methodology for finding similarity and co-reference of documents across languages. The similarity between the documents is identified according to the content of the whole document and co-referencing of documents is found by taking the named entities present in the document. Here we use Vector Space Model (VSM) for identifying both similarity and co-reference. This can be applied in cross-lingual search engines where users get documents of very similar content from different language documents.