Ontology based semantic similarity comparison of documents

Vladimir A. Oleshchuk, Asle Pedersen · 2004

In this paper, we consider ontologies as knowledge structures that specify terms, their properties and relations among them to enable knowledge extraction from texts. We represent ontologies using a graph-based model that reflect semantic relationship between concepts and apply them to text analysis and comparison. Instead of raw document comparison we compare document footprint enhanced with concepts from the ontology (using different enhancement algorithms). The result of this process may be that documents not similar prior to the enhancement become similar (semantically on some abstraction level) after the enhancement. This is because the enhancement process may introduce in the document footprint abstract concepts from the ontology. Using the ontology we can enhance the foot-prints by adding concepts that are not present in the original document. We may use synonyms for a horizontal expansion and broader terms/superclasses/types in a vertical expansion or both for that matter.

Read the paper · More papers on PaperTik