Document Retrieval and Clustering: from Principal Component Analysis to Self-aggregation Networks.
Chris H. Q. Ding · 2003
Ve first extend Hopfield networks to cluster- ing bipartite graphs (words-to-docmnent association) and show that the solution is the principal component analysis. . then generalize this via. the rain-max clustering principle into a self-aggregation networks which arc composed of scaled PCA components via Hcbb rule. Clustering amounts to an updating proc(ss where connections between different clusters are automatically suppressed while connections within same clusters arc enhanced. This framework combines dimension reduction with clustering via neural networks and PCA. Self-aggregation networks can also improve infbrmation retrieval performance. Applications are presented.