A GA-Based Approach to Fine-Tuning BERT for Hate Speech Detection

Kosisochukwu Judith Madukwe, Xiaoying Gao, Bing Xue · 2020

There exists a high level of variance in the finetuning of contextual word embedding models like Bi-directional Encoder Representation for Transformers (BERT). Such variance can be introduced by various factors, such as the BERT encoder layers, the fine-tuning architecture or the hyperparameter settings. Each architectural design choice leads to different results for a specific task. Hence, we are interested in reducing this variance given some settings. This study illustrates the use of a Genetic Algorithm (GA) to search, select and design a (near-) optimal fine-tuned BERT architecture for a hate speech detection task. We propose an appropriate encoding scheme for this task which represents the possible solutions in a way that supports a less time-consuming search for the global optima. Each encoding for a single solution represents both the BERT architecture and the fine-tuning architecture for a robust search. The automatic search provided by the GA, helps to reduce the time and cost involved in manual trial and error design methods. The experiments show that the resulting architectural design and hyperparameter settings are good choices for the hate speech detection task. Our method, although validated only on hate speech detection tasks, can easily be extended and generalized to other text classification tasks.

Read the paper · More papers on PaperTik