Multi-label Classification of Croatian Legal Documents Using EuroVoc Thesaurus

Frane Šarić, Bojana Dalbelo Bašić, Marie‐Francine Moens, Jan Šnajder · Lirias · 2014

The automatic indexing of legal documents can improve access to legislation. In this paper we describe the work on EuroVoc indexing of Croatian legislative documents. We focus on the machine learning aspect of the problem. First, we describe the manually indexed Croatian legislative documents collection, which we make freely available. Secondly, we describe the multi-label classification experiments on this collection. A challenge of EuroVoc indexing is class sparsity, and we discuss some strategies to address it. Our best model achieves 79.6% precision, 60.2% recall, and 68.6% F1-score.

Read the paper · More papers on PaperTik