Expanding Training Dataset with Class Hierarchy in Hierarchical Text Categorization
LI Baol · Beijing Daxue Xuebao. Zirankexueban · 2015
As the number of classes is quite large in a hierarchical text categorization problem, it usually costs much to obtain a training dataset of reasonable size and sample distribution. Several strategies are proposed and compared to generate new training samples from the class hierarchy in a hierarchical text classification problem. These solutions try to make full use of the class hierarchy(including class names, their descriptions if any, and relationships between them), and derive new pseudo training samples based on connotations and extensions of classes. Experiments on the dataset of the first large scale Chinese News Categorization at NLPCC 2014 show that the localized expanding strategy based on class extensions performs better. The proposed official system achieved Macro F1 0.8413 and 0.7139 at level 1 and level 2 respectively, which ranked the proposed system the second place among the 10 participating systems.