Offense Detection Using BERT and CNN
Meet Mehta, Dhruv Gada, Riddhi Sharma, Khushi Chavan, Pratik Kanani · 2022 IEEE 3rd Global Conference for Advancement in Technology (GCAT) · 2022
In today's unconventional world, the amount of user-generated content, such as articles, books, blogs, etc., is increasing exponentially. Some might find them offensive. As is well known, RNN and LSTM take a long time to analyze input since the words are processed one at a time; however, this happens all at once in BERT. In this paper, the suggested model uses BERT with CNN layers. These CNN layers help the model to improve its accuracy further. Hate Speech and Offensive Language is the dataset used. In this dataset there were three categories hate, offensive and neither, the accuracy with these was about 91%, since the paper focused on offense, the two categories were taken into consideration, offensive and not offensive. The accuracy of BERT without CNN was 85%, whereas BERT with CNN 2 categories was almost 95%.