Classification methods for documents with both fixed and free formats by PLSI model

Shigeichi Hirasawa, Wesley W. Chu · 2004

Related area: data base, information retrieval, knowledge acquisition, text mining, classification, clustering Abstract — Based on information retrieval model especially probabilistic latent semantic indexing (PLSI) model, we discuss methods for classification and clustering of a set of documents. A method for classification is presented and is demonstrated its good performance by applying to a set of benchmark documents with free format (text only). Then the classification method is modified to a clustering method and the clustering method is applied to partition experimental documents with fixed and free formats into two clusters, where the experimental documents are obtained from student questionnaires. Since the experimental documents are already categorized, the clustering method can be clearly evaluated its performance. The method has better performance compared to the conventional one based on the vector space model. The purpose of these questionnaires is to obtain useful knowledge for improvements in quality of education.

Read the paper · More papers on PaperTik