A Novel Semantic Approach to Document Collections

Andrea Addis, Manuela Angioni, Giuliano Armano, Roberto Demontis, Franco Tuveri, Eloisa Vargiu · 2008

ABSTRACT Available document collections are more and more required for supervised text categorization tasks. They typically are collections of documents classified by domain engineers. In this paper, we propose a semantic text categorization approach able to automatically create document collections in which documents are classified according to WordNet Domains taxonomy. Experiments have been performed by training a classifier with an automatic document collection and comparing results with those obtained by training the same classifier on a hand-made document collection. Experimental results point out that, on average, the performances of the automatic approach are quite similar to those obtained on a document collection classified by domain engineers. KEYWORDS Text Categorization, Document Collections, Intelligent Software Systems, Machine Learning. 1. INTRODUCTION Text categorization can be defined as the task of determining and assigning topical labels to content. The more the amount of available data (e.g., in digital libraries), the greater the need for high-performance text categorization algorithms. In particular, text categorization is a key technology in several information processing tasks, including controlled vocabulary indexing, routing and packaging of news and other text streams, content filtering, information security, help desk automation, and others. In the literature, many machine learning approaches have been proposed, both in the field of supervised [Sebastiani02] and unsupervised [Ghahramani04] learning. In particular, supervised approaches use only labeled data during the training phase. On the contrary, unsupervised approaches use unlabeled data, which may be easy to collect, but difficult to use. Semi-supervised learning in part resolves this problem by using large amount of unlabeled data, together with labeled data, to build better classifiers [Zhu05]. In this scenario, available text categorization document collections are more and more required. They are typically standard collections to which humans have assigned categories from a predefined set ([Lewis96], [Yang99], and [Lewis04]), so that researchers are able to test their algorithms in a controlled benchmarking context. Unfortunately, existing document collections suffer from one or more of the following drawbacks: (i) few documents, (ii) lack of the document full text, (iii) inconsistent or incomplete category assignment, (iv) peculiar textual properties, and (v) limited availability. Moreover, often researchers do not have documentation on how collections were produced, and on the nature of the underlying categories. To our knowledge, so far, only few attempts to automatically create document collections have been proposed [Ko00]. In particular, semantic approaches to text categorization have not been applied to this specific task. In this paper, we illustrate a method to create document collections by adopting a fully-automated semantic approach. Each text document is suitably labeled according to a predefined taxonomy of classes, namely WordNet Domains [Magnini00]. Experimental results point out that the proposed method allows to create reliable document collections.

Read the paper · More papers on PaperTik