A novel approach to sequence-of-documents focused text categorization using the concept of a degree of fuzzy set subsethood

Sławomir Zadrożny, Janusz Kacprzyk, Marek Gajewski · 2015

This work is meant as a step towards developing an effective and efficient procedure for a special type of the text categorization problem. A set of documents and a set of their categories are assumed. However, in addition to being assigned to a specific category, each document belongs to a certain sequence of documents, referred to as a case, comprising of documents from the same class. The problem considered is how to classify a document to a proper sequence of documents, or case, within a specified category. If each case is treated as a separate category, then the potential training datasets are rather small. We propose an algorithm which is based on a combination of two indicators characterizing a document to be classified: one reflecting its similarity to a case and one reflecting its similarity to a category. These indicators are based on the measure of subsethood of fuzzy sets. We study the effectiveness of the proposed algorithm for various combinations of weights of both indicators and subsethood measures employed.

Read the paper · More papers on PaperTik