Nonparametric methods for automatic classification of documents and transactions (abstract)

Amos O. Olagunju · 1990

The question of how to classify documents is a central problem in document retrieval. The classification problem can be stated as follows. There exists a large document collection, each of which contains a set of terms. How should the documents be clustered to allow the selection of index terms so that the collection can be searched to the maximal collective benefit of the retrieval system customers? Traditionally, transaction functionalities are manually scheduled into deferred and immediate queues for processing without any special consideration given to the interwoven functionalities invoked by the different user groups. The question of how to classify transactions for the concurrency controller in a distributed system is a major problem in transaction scheduling. The problem here is, how should transaction functionalities be scheduled for processing to satisfy the requirements of the different user groups? That is, how should transaction functionalities be organized on disk to minimize disk access time, in the hope of fulfilling the requirements of individual user groups?This paper presents nonparametric algorithms and heuristic for automatic classification of documents according to the similarity in their keywords; the words likely to be useful as index terms for document set. The normal approximation to the binomial distribution was explored as an index for automatic classification of documents and transactions. A nonparametric measure of association consistent with the Cramer statistic was used in the examination of similarities among documents. A nonparametric analysis of variance procedure was developed for comparing the profiles of term frequencies between documents or transaction functionalities invoked between users. The usefulness of the heuristic in the automatic classification of user groups according to the transaction functionalities that they invoke in a distributed system is discussed.

Read the paper · More papers on PaperTik