MultiFin: A Dataset for Multilingual Financial NLP

Rasmus Lundhus Jørgensen, O. Brandt, Mareike Hartmann, Xiang Dai, Christian Igel, Desmond Elliott · 2023

Financial information is generated and distributed across the world, resulting in a vast amount of domain-specific multilingual data.Multilingual models adapted to the financial domain would ease deployment when an organization needs to work with multiple languages on a regular basis.For the development and evaluation of such models, there is a need for multilingual financial language processing datasets.We describe MULTIFIN-a publicly available financial dataset consisting of real-world article headlines covering 15 languages across different writing systems and language families.The dataset consists of hierarchical label structure providing two classification tasks: multi-label and multiclass.We develop our annotation schema based on a real-world application and annotate our dataset using both 'label by native-speaker' and 'translate-then-label' approaches.The evaluation of several popular multilingual models, e.g., mBERT, XLM-R, and mT5, show that although decent accuracy can be achieved in high-resource languages, there is substantial room for improvement in low-resource languages.

Read the paper · More papers on PaperTik