Parallel clustering and classification of monolithic and non-monolithic document bases
Anthony S. Ruocco, Ophir Frieder · 1996
Cluster based approaches to information retrieval have been hampered by computationally intensive processes required for organizing the document base, and maintaining the document base in a dynamic environment. The high degree of computational requirements was conducive to the use of high performance, parallel processing computers. The hypothesis was that the commonly used single-pass clustering algorithm could be parallelized to overcome the time-prohibitive clustering process of massive document sets. This was verified by large reductions in clustering time while attaining near linear speed up in applications requiring a high level of similarity among documents in a cluster. The second focus was in using parallelized classification algorithms to maintain document set classification schemes in a dynamic environment. The hypothesis was that parallel processing could be used to reduce the large storage requirements (several hundred megabytes) of ancillary data required for document set classification. This was verified by having the algorithm produce data internally as required. The massive storage requirement was eliminated, and classification of large sets was performed in a non-prohibitive time frame while attaining reasonable speed up. The third area of focus centered on the document base model itself. All current algorithm and retrieval measures are based on all documents being collocated in a single document base. The premise was that the current monolithic document base model does not adequately reflect the emerging environment of distributed, and independent information systems. The hypothesis was that queries conducted at the independent system level would eliminate the need for building and maintaining large, unwieldy, monolithic structures. This was verified by treating partitions of the document base as independent entities which consistently resulted in the formation of tighter clusters at, what would be, the independent system level.