WordNet-based text document clustering
Julian Sedding, Dimitar Kazakov · 2004
Text document clustering can greatly simplify browsing large collections of documents by reorganizing them into a smaller number of manageable clusters.Algorithms to solve this task exist; however, the algorithms are only as good as the data they work on.Problems include ambiguity and synonymy, the former allowing for erroneous groupings and the latter causing similarities between documents to go unnoticed.In this research, naïve, syntax-based disambiguation is attempted by assigning each word a part-of-speech tag and by enriching the 'bag-ofwords' data representation often used for document clustering with synonyms and hypernyms from WordNet.