Distributional semantics study using the co-occurrence computed from collaborative resources and WordNet
Mohamed Ben Aouicha, Mohamed Ali Hadj Taieb, Sameh Beyaoui · 2016
Quantifying the semantic relation between words is a key element in several applications including the treatments at the meaning level. A great variety of approaches are proposed in order to quantify the semantic proximity between concepts or words. These approaches exploit computational models including the hierarchical and textual information of the semantic resources. Among these models, the distributional approaches quantify the semantic relations based on the co-occurrence information according to the target words. In this paper, we study the distributional semantics of three resources: the collaborative resources Wiktionary and Wikipedia, and the thesaurus WordNet through the word relatedness task. We exploit the glosses of WordNet and Wiktionary as a corpus formed by short and precise words, and the contents of Wikipedia articles. The experiments are performed using the known measures PMI and cosine, and a list of known benchmarks in semantic relatedness task. The results show that a small corpus formed by well formed sentences can lead to good correlations but limited coverage capacity. Despite the improvement in coverage capacity using Wikipedia, the correlations between human judgments and computed values do not follow the same enhancement degree.