Information Content Methods for Semantic Similarity: An Experimental Assessment

Antonio De Nicola, Anna Formica, Ida Mele, Francesco Taglino · IEEE Access · 2025

This paper compares some of the most representative methods for calculating theinformation content(IC) of concepts within a taxonomy. These methods fall into two categories: extensional methods, which rely on external resources such as text corpora, and intensional methods, which depend solely on the taxonomy’s internal structure. To evaluate these methods, we integrated each one intoSemSimp, a technique used to measure the similarity between resource annotations (i.e., sets of concepts from a taxonomy). We show how the adoption of each IC method affects the overall performance ofSemSimpand the resulting similarity scores between annotated resources. To assess the results, we used three metrics: (i) the degree of confidence (a statistical measure); (ii) the Pearson correlation (based on expert judgments); (iii) the harmonic mean of the two. Although previous studies suggest that extensional methods outperform intensional ones, our findings reveal that both categories of methods can achieve strong performance in semantic similarity tasks. Specifically, the best-performing methods achieved an average degree of confidence of 0.86, an average Pearson correlation of 0.69, and their harmonic mean of 0.75.

Read the paper · More papers on PaperTik