Similar document detection using self-organizing maps
Anssi Lensu, Pasi Koikkalainen · 2003
This paper describes how similar free-form textual documents can be matched using the self-organizing maps (SOMs). The analysis chain is made of three parts: first, similar words are located using an alphabet occurrence coding and SOM; second, three-word contexts are clustered using codes obtained from the word SOM to build a context map; and third, whole documents are clustered using codes from the context SOM. Although this work is inspired by the WEBSOM method, it is quite different since our goal was to build a fast system, which is tolerant to the special features of different languages.