A Comparative Study of Deep Neural Network Models on Multi-Label Text Classification in Finance
Macedo Maia, Juliano Efson Sales, André Azul Freitas, Siegfried Handschuh, Markus Endres · 2021
Multi-Label Text Classification (MLTC) is a well-known NLP task that allows the classification of texts into multiple categories indicating their most relevant domains. However, training model tasks on texts from web user deal with redundancy or ambiguity of linguistic information. In this work, we propose a comparative study about different neural network models for a multi-label text categorisation task in finance domain. Our main contribution consists of presenting a new annotated dataset that contains~26k posts from users associated to finance categories. To build that dataset, we defined 10 specific-domain categories that cover financial texts. To serve as a baseline, we present a comparative study analysing both the performance and training time of different learning models for the task of multilabel text categorisation on the new dataset. The results show that transformer-based language models outperformed RNN-based neural networks in all scenarios in terms of precision. However, transformers took much more time than RNN models to train an epoch model.