SemSim: Resources for Normalized Semantic Similarity Computation Using Lexical Networks

Elias Iosif, Alexandros Potamianos · 2012

We investigate the creation of corpora from web-harvested data following a scalable approach that has linear query complexity.Individual web queries are posed for a lexicon that includes thousands of nouns and the retrieved data are aggregated.A lexical network is constructed, in which the lexicon nouns are linked according to their context-based similarity.We introduce the notion of semantic neighborhoods, which are exploited for the computation of semantic similarity.Two types of normalization are proposed and evaluated on the semantic tasks of: (i) similarity judgement, and (ii) noun categorization and taxonomy creation.The created corpus along with a set of tools and noun similarities are made publicly available.

Read the paper · More papers on PaperTik