A Novel Cost-efficient Use of BERT Embeddings in 8-way Emotion Classification on a Hungarian Media Corpus
György Márk Kis, Orsolya Ring, Miklós Sebők · 2022
BERT is a state-of-the-art open-sourced NLP solution for a wide range of tasks, but even basic uses need extensive GPU resources, technical expertise, and a large and good quality corpus with wide domain coverage. This paper presents an alternative, cost-efficient approach for using the strength of BERT’s tokenizer and contextual embeddings for virtually free, with only nominal Python and Machine Learning expertise. This method of using traditional ML-modeling enhanced by BERT’s contextual embeddings extracted from the hidden layers gives worse but comparable results to fine-tuned full BERT models. On our novel, manually validated dataset of Hungarian online media texts, seven emotions, and a neutral class were classified with a 0.62 weighted F1-score. While two categories were underperforming due to small sample representation and size, the other six produced results that sit squarely between classical frequency-based methods (0.44-0.47) and the newer, fine-tuned BERT models (0.71-0.73). We outline several routes for improvements and conclude that while ’doodling’ around with BERT is a large undertaking for many researchers, our approach provides good results with a fraction of the effort needed. The gap between our approach and a fine-tuned BERT significantly narrows down on corpuses necessitating parameter degradation due to insufficient memory.