Unsupervised Multi-Label Document Classification for Large Taxonomies Using Word Embeddings
Stefan Hirschmeier, Johannes Werner Melsbach, Detlef Schoder, Sven Stahlmann · 2019
More and more businesses are in need for metadata for their documents. However, automatic generation for metadata is not easy, as for supervised document classification, a significant amount of labelled training data is needed, which is not always present in the desired amount or quality. Often, documents need to be tagged with a predefined set of company specific keywords that are organized in a taxonomy. We present an unsupervised approach to perform multi-label document classification for large taxonomies using word embeddings and evaluate it with a dataset of a public broadcaster. We point out strengths of the approach compared to supervised classification and statistical approaches like tf-idf.