Named entities as privileged information for hierarchical text clustering
Roberta Akemi Sinoara, Camila Vaccari Sundermann, Ricardo Marcondes Marcacini, Marcos Aurélio Domingues, Solange Oliveira Rezende · 2014
Text clustering is a text mining task which is often used to aid the organization, knowledge extraction, and exploratory search of text collections. Nowadays, the automatic text clustering becomes essential as the volume and variety of digital text documents increase, either in social networks and the Web or inside organizations. This paper explores the use of named entities as privileged information in a hierarchical clustering process, so as to improve clusters quality and interpretation. We carried out an experimental evaluation on three text collections (one written in Portuguese and two written in English) and the results show that named entities can be applied as privileged information to power clustering solution in dynamic text collection scenarios.