Automatic subject heading assignment for online government publications using a semi‐supervised machine learning approach

Xiao Hu, Larry S. Jackson, Sai Deng, Jing Zhang · Proceedings of the American Society for Information Science and Technology · 2005

Abstract As the dramatic expansion of online publications continues, state libraries urgently need effective tools to organize and archive the huge number of government documents published online. Automatic text categorization techniques can be applied to classify documents approximately, given a sufficient number of labeled training examples. However, obtaining training labels is very expensive, requiring a lot of manual labor. We present a real world online government information preservation project (PEP) in the State of Illinois, and a semi‐supervised machine learning approach, an Expectation‐Maximization (EM) algorithm‐based text classifier, which is applied to automatically assign subject headings to documents harvested in the PEP project. The EM classifier makes use of easily obtained unlabeled documents and thus reduces the demand for labeled training examples. This paper describes both the context and the procedure of such an application. Experiment results are reported and other alternative approaches are also discussed.

Read the paper · More papers on PaperTik