Multilabel Clustering Analysis of the Croatian-English Parallel Corpus Based on Latent Dirichlet Allocation Algorithm

Erzsébet Tóth, Zoltán Gál · 2023

A parallel corpus of Croatian EU legislative documents translated automatically to English over 28 years with a year of creation and hierarchical classifier tags including descriptors, document types, and fields considered as meta information assigned to each text. Only two third part of around 1.5 thousand texts have all the fields completed, accomplishing the required manual work too time-consuming for human administration. Similar incompleteness of legal texts may appear in official legal sites operated as regular service provisioning databases. In this paper we proposed an artificial cognitive and multilabel classification method to automatically find the necessary tags for the corpus with just a tiny fraction of the manual tagging time. The Latent Dirichlet Allocation algorithm assigns field values or tags to incompletely labelled documents. The dependence of the quantitative linguistics properties was presented in the function of the type and specialty of pre-processing tasks. We successfully applied this algorithm built on no error correcting optimising codes to predict a mixture of topic probabilities of these legal texts on the basis of Hamming distance of the binary feature vectors created using the legal fields of the EUROVOC multilingual thesaurus.

Read the paper · More papers on PaperTik