Extraction of Folksonomies from Noisy Texts

Wim De Smet, Marie‐Francine Moens · 2007

We built a system for the automatic creation of a text-based topic hierarchy, meant to be used in a geographically defined community. This poses two main problems. First, the ap-pearance of both standard language and a community-related dialect, demanding that dialect words should be as much as possible corrected to standard words, and second, the auto-matic hierarchic clustering of texts by their topic. The problem of correcting dialect words is dealt with by performing a nearest neighbor search over a dynamic set of known words, using a set of transition rules from dialect to standard words, which are learned from a pair-wise lexicon. We tackle the clustering problem by implementing a hierar-chical co-clustering algorithm that automatically generates a topic hierarchy of the collection and simultaneously groups documents and words into clusters.

Read the paper · More papers on PaperTik