Dataless Text Classification with Pseudo Topic Representation
Rong Yan, Qi Chen, Guanglai Gao · 2020
As for an automatic text classification approach, a large body of research on latent-topic based Dataless Text Classification (DTC) has been emerged in recent years. Perusing the candidate seed words or guaranteeing the quality of the category-topics is the core mission of this approach. However, few previous studies consider the quality of specific category-topics at the collection level instead at the document level, because not all topics are equally coherent or category sparsity. In this paper, we focus on alleviating the dilemma for the seed words selection problem in DTC by using pseudo text understanding. Differently from the existing latent-topic based DTC approach, we propose an unsupervised method named Pseudo Document Labeled Classification (PDLC). It extracts the most representative word list to capture the best latent semantic category-topic description. Experimental results indicate that our PDLC scheme achieves better classification accuracy without any labeled data or external resource.