Experiment with a hierarchical text categorization method on the WIPO-alpha patent collection

Domonkos Tikk, György Biró · 2004

Text categorization is the classification to assign a text document to an appropriate category in a predefined set of categories. We focus on the special case when categories are organized in hierarchy. We present a new approach on this recently emerged subfield of text categorization. The algorithm applies an iterative learning module that allow of gradually creating a classifier by trial-and-error-like method. We present a software that has been developed on the basis of the algorithm to illustrate the capability of the algorithm on large data collection. We experimented on the very large benchmark collection, on the WIPO-alpha (World Intellectual Property Organization, Geneva, Switzerland, 2002) English patent database that consists of about 75000 XML documents distributed over 5000 categories. Our software is able to index the corpus quickly and creates a classifier in a few iteration cycles. We present the results achieved by the classifier w.r.t. various test setting

Read the paper · More papers on PaperTik