A study of document filtering using the subspace method of pattern recognition
Tsutomu Matsunaga · Systems and Computers in Japan · 2000
One of the typical filtering approaches, in which no constraint is imposed on the document structure, is the method based on the vector space model. In this method, the words composing the document are assigned to the elements of the vector, and the document is selected statistically by the similarity in the constructed feature vector space. Hitherto, the documents have been ranked in the order of high similarity, and the one-dimensional filtering is considered, based on the specified number of documents in the upper rankings or a preset threshold. A problem in this filtering is that only the necessary information is selected following the order, and noise elimination to reject the unnecessary information is not directly considered. From such a viewpoint, this study takes the approach of considering the user's interest, both positive and negative. The filtering is considered as the corresponding two-category problem, and a filtering method is proposed based on the concept of the pattern classification. This method uses the subspace method, which is known as a method for pattern recognition. The proposed method consists of highly precise filtering introducing the co-occurrence relation among words, and realizes the representation and updating of the interest items by a single mechanism. The effectiveness of the proposed method is demonstrated by experiments using newspaper articles. © 1999 Scripta Technica, Syst Comp Jpn, 31(1): 48–58, 2000