Using Kohonen maps to determine document similarity

Jennifer Farkas · Conference of the Centre for Advanced Studies on Collaborative Research · 1994

In this paper we present some experimental results on the classification of natural language documents using Kohonen's self-organizing-map neural network paradigm. We discuss, in particular, how the classification accuracy can be improved if the standard keyword representation of documents is enhanced by including specific weights, thesaurally-defined relations among keywords, and additional synonyms for keywords. We sketch the main features of a prototype of an automatic document classification system which is capable of classifying full-text documents relative to a controlled domain-specific vocabulary and thesaural relations. The described results extend earlier work on the use of neural networks for clustering semantically similar documents.

Read the paper · More papers on PaperTik