Self-organizing maps of massive document collections
Teuvo Kohonen · 2000
Huge document collections can be organized according to textual similarities by the self-organizing map (SOM) algorithm, when statistical representations of the textual contents are used as the feature vectors of the documents. In a practical experiment we mapped 6,840,568 patent abstracts onto a 1,002,240-node SOM. For the feature vectors we selected 500-dimensional random projections of the weighted word histograms.