WordNet-based text document clustering

Julian Sedding, Dimitar Kazakov · 2004

Text document clustering can greatly simplify browsing large collections of documents by reorganizing them into a smaller number of manageable clusters.Algorithms to solve this task exist; however, the algorithms are only as good as the data they work on.Problems include ambiguity and synonymy, the former allowing for erroneous groupings and the latter causing similarities between documents to go unnoticed.In this research, naïve, syntax-based disambiguation is attempted by assigning each word a part-of-speech tag and by enriching the 'bag-ofwords' data representation often used for document clustering with synonyms and hypernyms from WordNet.

Read the paper · More papers on PaperTik