Tag Similarity in Folksonomies
Hatem Mousselly-Sergieh, Elöd Egyed-Zsigmond, Gabriele Gianini, Mario Döller, Harald Kosch, Jean-Marie Pinon · 2013
ABSTRACT. Folksonomies- collections of user-contributed tags, proved to be efficient in reducing the inherent semantic gap. However, user tags are noisy; thus, they need to be processed before they can be used by further applications. In this paper, we propose an approach for bootstrapping semantics from folksonomy tags. Our goal is to automatically identify semantically related tags. The approach is based on creating probability distribution for each tag based on co-occurrence statistics. Subsequently, the similarity between two tags is determined by the distance between their corresponding probability distributions. For this purpose, we propose an extension for the well-known Jensen-Shannon Divergence. We compared our approach to a widely used method for identifying similar tags based on the cosine measure. The evaluation shows promising results and emphasizes the advantage of our approach. RÉSUMÉ. Les folksonomies sont des collections d’annotations créés de manière collaborative par plusieurs utilisateurs. Afin d’améliorer la recherche d’information à l’aide de folksonomies, leur classification permet d’apporter des solutions à certains des problèmes inhérentes au caractère libre et collaboratif des folksonomies. Ces problèmes sont les fautes de frappe, des séparations-collage d’expressions, le multilinguisme, etc. Dans ce papier nous proposons une nouvelle mé-